Virtual interaction method and device based on voice equipment and storage medium
By determining the interaction space and influence level of voice devices, combining voice and video data, a virtual interaction mechanism is built, and the problem of insufficient interaction accuracy in the existing technology is solved, and more accurate and smooth voice device interaction is achieved.
Patent Information
- Application Number
- CN202510348162.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing voice devices do not fully consider the interaction influence level in the interactive space, resulting in insufficient accuracy of the virtual interaction mechanism.
By determining the interaction space, interaction influence level, interaction path and service life of voice devices, combining voice and video data, a virtual interaction mechanism is built to optimize the interaction strategy of voice devices.
It improves the interaction accuracy and smoothness of voice devices in various scenarios, and realizes refined interactive control of users and the environment.
Smart Images

Figure CN120260559A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual interaction based on voice devices, and particularly relates to a virtual interaction method, device and storage medium based on voice devices. Background Art
[0002] With the development of technology, voice devices have been gradually applied to life and can interact with users in various ways. At this time, voice devices cover various interactive devices such as smart speakers and smart robots. In the prior art, voice devices collect voice information of users and provide corresponding voice feedback based on the voice information, which is only limited to a single interaction space and does not consider the influence of the interaction impact level of the interaction space, thus affecting the accuracy of the virtual interaction mechanism of voice devices. Summary of the Invention
[0003] Based on this, it is necessary to provide a virtual interaction method, device and storage medium based on voice devices for the above technical problems.
[0004] A virtual interaction method based on voice devices includes:
[0005] Determining the interaction space of the voice device based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range;
[0006] Determining the interaction impact level of the interaction space according to the environmental detection of the interaction space of the voice device;
[0007] Determining the interaction path of the voice device based on the location of the voice device and the location of the user;
[0008] Determining the virtual interaction mechanism of the voice device according to the interaction path of the voice device, the interaction impact level of the interaction space, and the service life of the voice device;
[0009] In the virtual interaction mechanism of the voice device, determining a virtual interaction data set based on the video picture and voice data collected by the voice device; outputting the virtual interaction content of the user and the virtual interaction content of the environment according to the recognition of the virtual interaction data set;
[0010] Determining a virtual interaction list according to the virtual interaction content of the user, the interaction part of the voice device, the priority matching table of interaction actions, and the work schedule of the user, and at the same time, determining the expression factor of each interaction item in the virtual interaction list based on the virtual interaction content of the environment and the interaction part of the voice device.
[0011] A virtual interaction device based on voice devices is applied to the above virtual interaction method based on voice devices. The virtual interaction device based on voice devices includes:
[0012] An interaction space module, configured to determine the interaction space of a voice device based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range;
[0013] An interaction impact level module, configured to determine the interaction impact level of the interaction space according to the environmental detection of the interaction space of the voice device;
[0014] An interaction path module, configured to determine the interaction path of the voice device based on the location of the voice device and the location of the user;
[0015] A virtual interaction mechanism module, configured to determine the virtual interaction mechanism of the voice device according to the interaction path of the voice device, the interaction impact level of the interaction space, and the service life of the voice device;
[0016] A virtual interaction content module, configured to determine a virtual interaction data set based on the video picture and voice data collected by the voice device in the virtual interaction mechanism of the voice device; output the virtual interaction content of the user and the virtual interaction content of the environment according to the recognition of the virtual interaction data set;
[0017] A virtual interaction list module, configured to determine a virtual interaction list according to the virtual interaction content of the user, the interaction part of the voice device, the priority matching table of interaction actions, and the work schedule of the user. At the same time, determine the expression factors of each interaction item in the virtual interaction list based on the virtual interaction content of the environment and the interaction part of the voice device.
[0018] A storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above virtual interaction method based on a voice device is implemented;
[0019] In the embodiment of the present invention, through the method in the embodiment of the present invention, the interaction space of the voice device is determined based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range; the interaction impact level of the interaction space is determined according to the environmental detection of the interaction space of the voice device; the interaction path of the voice device is determined based on the location of the voice device and the location of the user; the virtual interaction mechanism of the voice device is determined according to the interaction path of the voice device, the interaction impact level of the interaction space, and the service life of the voice device, which comprehensively considers the interaction path of the voice device, the interaction impact level of the interaction space, and the service life of the voice device, improves the accuracy of the virtual interaction mechanism of the voice device, and enables the voice device to perform smooth online interaction with the user in various scenarios.
[0020] Further, in the virtual interaction mechanism of the voice device, a virtual interaction data set is determined based on the video images and voice data collected by the voice device; the virtual interaction content of the user and the virtual interaction content of the environment are output according to the recognition of the virtual interaction data set, and the virtual interaction content of the user and the virtual interaction content of the environment are introduced to facilitate the refined control of the virtual interaction content of the user and the virtual interaction content of the environment.
[0021] Therefore, a virtual interaction list is determined according to the virtual interaction content of the user, the interaction part of the voice device, the priority matching table of interaction actions, and the work schedule of the user. At the same time, the expression factors of each interaction item in the virtual interaction list are determined based on the virtual interaction content of the environment and the interaction part of the voice device, ensuring the accuracy of the virtual interaction list and introducing expression factors into the virtual interaction list, realizing the phased interaction actions of the voice device and the corresponding expression control. Brief Description of the Drawings
[0022] Figure 1 It is a schematic diagram of the application scenario of the virtual interaction method based on a voice device in an embodiment;
[0023] Figure 2 It is a schematic flowchart of the virtual interaction method based on a voice device in an embodiment;
[0024] Figure 3 It is a structural block diagram of the virtual interaction device based on a voice device in an embodiment;
[0025] Figure 4 It is the internal structure diagram of an electronic device in an embodiment. Detailed Embodiments
[0026] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0027] Embodiment 1
[0028] The virtual interaction method based on a voice device provided by the present application can be applied to an application environment as Figure 1 shown. Among them, the computer 102 communicates with the server 104 through the network. Among them, the terminal 102 can be, but is not limited to, various personal computers, servers, and voice devices, and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0029] Embodiment 2
[0030] In this embodiment, please refer toFigures 2 to 4 , a virtual interaction method based on a voice device, which is applied to a virtual interaction scenario based on a voice device; the virtual interaction method based on a voice device includes:
[0031] Step S11: Determine the interaction space of the voice device based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range.
[0032] Step S12: Determine the interaction influence level of the interaction space according to the environmental detection of the interaction space of the voice device.
[0033] Step S13: Determine the interaction path of the voice device based on the location of the voice device and the location of the user.
[0034] Step S14: Determine the virtual interaction mechanism of the voice device according to the interaction path of the voice device, the interaction influence level of the interaction space, and the service life of the voice device.
[0035] Step S15: In the virtual interaction mechanism of the voice device, determine a virtual interaction data set based on the video picture and voice data collected by the voice device; output the virtual interaction content of the user and the virtual interaction content of the environment according to the recognition of the virtual interaction data set.
[0036] Step S16: Determine a virtual interaction list according to the virtual interaction content of the user, the interaction part of the voice device, the priority matching table of interaction actions, and the work schedule of the user. At the same time, determine the expression factors of each interaction item in the virtual interaction list based on the virtual interaction content of the environment and the interaction part of the voice device.
[0037] In step S11, determine the interaction space of the voice device based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range.
[0038] In the specific implementation process of the present invention, the specific steps may be:
[0039] S111: In the indoor space, determine the voice device based on the traversal of the indoor space, and mark the location of the voice device.
[0040] S112: Determine the voice collection module according to the online detection of the voice device, and determine the theoretical voice collection range of the voice device according to the traversal of the voice collection module; determine the actual voice collection range of the voice device according to the test of multiple voice collection nodes in the theoretical voice collection range of the voice device.
[0041] S113: Determine the video capture module based on the online detection of the voice device, and determine the theoretical video capture range of the voice device according to the traversal of the video capture module; determine the actual video capture range of the voice device according to the tests of multiple video capture nodes in the theoretical video capture range of the voice device;
[0042] S114: Interact the location of the voice device, the actual voice capture range of the voice device, and the actual video capture range, and output multiple spatial features according to the interaction of the location of the voice device, the actual voice capture range of the voice device, and the actual video capture range;
[0043] S115: Determine the interaction space of the voice device according to multiple spatial features, the location of the voice device, and the location of the user.
[0044] In the embodiments of the present application, in an indoor space, the voice device is determined based on the traversal of the indoor space, and the location of the voice device is marked. The application scenario of the indoor space is introduced, and the indoor space is traversed at the spatial level, so as to locate the voice device in the indoor space and further control the voice device subsequently.
[0045] Optionally, use tools such as CAD drawings, 3D modeling software, or laser scanners to create a 3D model of the indoor space. This 3D model of the indoor space should include information such as the shape, size, and furniture layout of the room. Create a corresponding traversal path based on the 3D model of the indoor space, and traverse along the traversal path to facilitate collecting the Bluetooth signal strength output based on the voice device during the traversal, and perform preliminary positioning according to the relationship between the signal strength and the distance, and use TDOA or beamforming technology to accurately calculate the location of the voice device.
[0046] Furthermore, determine the voice capture module based on the online detection of the voice device, and determine the theoretical voice capture range of the voice device according to the traversal of the voice capture module; determine the actual voice capture range of the voice device according to the tests of multiple voice capture nodes in the theoretical voice capture range of the voice device, ensuring the accuracy of the actual voice capture range of the voice device and realizing the tests of multiple voice capture nodes in the voice capture range of the theoretical voice device.
[0047] At this time, it is recognized through online detection that the voice capture module of the voice device is performing a voice capture task, so as to introduce the voice capture module. This usually involves analyzing the test signals sent by the device and checking the data returned by the device.
[0048] Meanwhile, each voice acquisition module is tested one by one to evaluate its performance, and the theoretical voice acquisition range of the voice device is calculated based on the technical specifications and performance test results of the voice acquisition module. Further, multiple nodes are set within the theoretical voice acquisition range for actual testing of the voice device's acquisition ability. A test voice signal is sent at each node, and it is observed whether the voice device can accurately acquire and recognize these signals. The actual voice acquisition range of the voice device is determined according to the test results, which is usually a smaller area than the theoretical range because there are various interference factors (such as noise, obstacles, etc.) in the actual environment.
[0049] Optionally, assume there is a smart home system with two voice assistant devices, named "Voice Assistant A" and "Voice Assistant B" respectively. The system determines through online detection that both devices are online and active, so they are both recognized as the current voice acquisition modules. The system conducts a traversal test on these two devices, including sending test voice signals and measuring the accuracy and integrity of the returned data. Based on the test results and the microphone sensitivity specifications of the devices, the system calculates that the theoretical voice acquisition range of "Voice Assistant A" is a circular area with a radius of 5 meters, and the theoretical voice acquisition range of "Voice Assistant B" is a circular area with a radius of 4 meters. Then, the system sets multiple voice acquisition nodes within the theoretical voice acquisition range of each device and conducts actual tests. It is found that "Voice Assistant A" can accurately acquire and recognize test signals within 4.5 meters during actual testing, while "Voice Assistant B" can only accurately acquire and recognize test signals within 3.5 meters. Therefore, the system determines that the actual voice acquisition range of "Voice Assistant A" is a circular area with a radius of 4.5 meters, and the actual voice acquisition range of "Voice Assistant B" is a circular area with a radius of 3.5 meters.
[0050] Further, the video acquisition module is determined based on the online detection of the voice device, and the theoretical video acquisition range of the voice device is determined according to the traversal of the video acquisition module; the actual video acquisition range of the voice device is determined according to the tests of multiple video acquisition nodes within the theoretical video acquisition range of the voice device.
[0051] At this time, through online detection, it is recognized that the video acquisition module of the voice device is performing a video acquisition task, so as to introduce the video acquisition module, which usually involves analyzing the test signals sent by the device and checking the data returned by the device.
[0052] Meanwhile, each video acquisition module is tested one by one to evaluate its performance. Based on the technical specifications and performance test results of the video acquisition module, the theoretical video acquisition range of the voice device is calculated. Further, multiple nodes are set within the theoretical video acquisition range for actually testing the acquisition ability of the voice device. A test voice signal is sent at each node, and it is observed whether the voice device can accurately acquire and recognize these signals. According to the test results, the actual video acquisition range of the voice device is determined, which is usually a smaller area than the theoretical range because there are various interference factors (such as noise, obstacles, etc.) in the actual environment.
[0053] Optionally, in the video acquisition module, based on parameters such as the focal length and viewing angle of the camera, the theoretical video acquisition range is calculated. For example, if the focal length of the camera is 5 mm and the viewing angle is 70 degrees, the maximum width and height that can be acquired at a certain distance can be calculated. Multiple test nodes are set within the theoretical video acquisition range, such as placing a marker every one meter. The intelligent voice device and the video acquisition module are started for acquisition, and the acquisition situation of each node is checked. By analyzing the test results, it is found that the acquisition effect of some edge nodes is not good (such as blurred images, markers being blocked, etc.), so the actual video acquisition range is determined to be slightly smaller than the theoretical range. For example, the actual acquisition range is an area within 3 meters in front of the camera and with a width corresponding to the viewing angle of the camera.
[0054] Therefore, the location of the voice device, the actual voice acquisition range of the voice device, and the actual video acquisition range are interacted, and multiple spatial features are output according to the interaction of the location of the voice device, the actual voice acquisition range of the voice device, and the actual video acquisition range. The interaction space of the voice device is determined based on the multiple spatial features, the location of the voice device, and the location of the user.
[0055] At this time, the location information of the voice device (such as coordinates, room number, etc.), the actual voice acquisition range (such as a circular or elliptical area centered on the device), and the actual video acquisition range (which can also be the definition of shape and size) are collected, and these data are input into the first interaction module, which is used to process these data and generate spatial features. The interaction processing involves spatial superposition, range comparison, position relationship analysis, etc. Based on the results of the interaction processing, multiple spatial features are output, and these features include the overlapping area of the voice and video acquisition ranges, the relative position relationship between the device and the user (such as distance, direction), the boundary conditions of the acquisition range, etc.
[0056] Furthermore, conduct in-depth analysis on the multiple spatial features output, understand how they affect the interaction between the voice device and the user, evaluate the interactivity between the user and the voice device based on the location of the voice device, the location of the user, and spatial features (such as overlapping areas, distance, and direction), and determine one or more interaction spaces according to the results of the location relationship evaluation. These spaces represent the areas where the user can effectively interact with the voice device, and also consider the impact of factors such as environmental noise and line-of-sight occlusion on the interaction.
[0057] Optionally, collect the exact location of the intelligent robot (such as the coordinates in the corner of the room), and collect the actual voice collection range of the intelligent robot (such as a circular area with a radius of 3 meters centered on the speaker) and the actual video collection range of the camera (such as an elliptical area with a 45-degree view centered on the camera). Input these data into the first interaction module, which generates spatial features based on the GIS algorithm, including the overlapping area of the voice and video collection ranges, the relative distance and direction between the speaker and the user, etc. Analyze the generated spatial features and find that the user has the highest interactivity with the intelligent robot within the overlapping area. Consider factors such as the environmental noise level (such as background music in the room) and camera line-of-sight occlusion (such as furniture layout), and fine-tune the interaction space. Finally, determine an interaction space, which is an elliptical area approximately 2 meters in front of the intelligent robot and is also within the video collection range of the camera.
[0058] In step S12, determine the interaction influence level of the interaction space according to the environmental detection of the interaction space of the voice device;
[0059] In the specific implementation process of the present invention, the specific steps can be:
[0060] S121: Obtain the interaction space of the voice device;
[0061] S122: Determine multiple spatial nodes based on the interaction space of the voice device and multiple detections of the voice device;
[0062] S123: Determine multiple environmental regions according to the multiple spatial nodes and the location of the voice device;
[0063] S124: Output multiple environmental parameters based on the environmental detection of the multiple environmental regions;
[0064] S125: Conduct multiple interactions on the multiple environmental parameters, the spatial positions of the multiple environmental regions, and the location of the voice device;
[0065] S126: Determine the first interaction influence parameter according to the multiple environmental parameters and the spatial positions of the multiple environmental regions, and determine the second interaction influence parameter according to the multiple environmental parameters and the location of the voice device;
[0066] S127: Determine the interaction influence level of the interaction space based on the first interaction influence parameter, the second interaction influence parameter, and the size of the interaction space.
[0067] In the embodiments of the present application, obtain the interaction space of the voice device; determine multiple spatial nodes based on the interaction space of the voice device and multiple detections of the voice device; determine multiple environmental regions according to the multiple spatial nodes and the location of the voice device, taking into account the overall situation of the multiple spatial nodes and the location of the voice device, and achieving precise control of the multiple environmental regions.
[0068] Specifically, obtaining the interaction space of the voice device introduces a clear interaction space of the voice device. The interaction space of the voice device is a physical area where users can obtain the best interaction experience when interacting with the voice device within this area. The determination of the interaction space is based on various factors, such as the voice collection range of the device, the video collection range, the activity range of the user, environmental factors (such as noise level, echo, etc.), and the relative position relationship between the device and the user.
[0069] After obtaining the interaction space, based on the multiple detection capabilities of the voice device (such as sound localization, direction perception, environmental noise analysis, etc.), determine multiple spatial nodes within the interaction space. These spatial nodes are key points of sound propagation, areas of sound reflection or interference, or key points of the camera's line of sight, etc. The purpose of determining these nodes is to better understand the acoustic characteristics, light conditions, and line of sight occlusion situations within the interaction space, so as to provide a basis for subsequent interaction optimization.
[0070] After determining the spatial nodes, the system can divide the interaction space into multiple smaller environmental regions according to these spatial nodes and the location of the voice device. These environmental regions have different acoustic characteristics (such as noise level, echo intensity), light conditions (such as light intensity, light direction), or line of sight occlusion situations (such as furniture layout, wall occlusion). By dividing the interaction space into multiple environmental regions, the system can more precisely analyze the interaction conditions within each environmental region and take corresponding optimization measures.
[0071] Optionally, after determining the interaction space, the system uses the sound localization ability of the intelligent voice assistant to determine several key space nodes, including the boundary points of the interaction space, areas of sound reflection or interference (such as near walls and large office furniture), and key points of the camera's line of sight (such as ensuring that the camera can capture most of the area within the interaction space). Based on the determined space nodes and the location of the intelligent voice assistant, the system divides the interaction space into three environmental areas. The first area is the core area of the interaction space, which is in front of the intelligent voice assistant and relatively close, with the best acoustic characteristics and line-of-sight conditions. The second area is the edge area of the interaction space, which is around the core area and is affected by some sound reflections or interferences. The third area is the external area of the interaction space. Although it is within the scope of the interaction space, due to factors such as being far away or being blocked by furniture, the interaction effect is not good.
[0072] Furthermore, by introducing the environmental detection of multiple environmental areas and outputting multiple environmental parameters based on it, the accuracy of multiple environmental parameters is ensured, so as to present the environmental conditions of multiple environmental areas through multiple environmental parameters.
[0073] At this time, environmental detection is carried out one by one on the previously determined multiple environmental areas, aiming to obtain specific environmental parameters within each area. These environmental parameters can reflect the acoustic, light, temperature, humidity and other characteristics of the area, which are crucial for subsequent analysis of the interaction effect and optimization of device configuration.
[0074] Furthermore, according to the application scenario and specific requirements, the environmental parameters to be detected are determined. For example, in the intelligent voice interaction scenario, parameters such as noise level, echo intensity, light intensity, temperature, and humidity need to be detected. After determining the parameters to be detected, the system needs to arrange corresponding detection devices in each environmental area.
[0075] For example, a noise meter can be used to measure the noise level, an echo tester can be used to evaluate the echo intensity, and a light sensor can be used to measure the light intensity, etc.
[0076] At the same time, after deploying the detection devices, the system starts to perform environmental detection, which includes measuring at different time periods and different environmental conditions (such as day and night, sunny and cloudy) to obtain more comprehensive and accurate data. At the same time, during the environmental detection process, the system will collect a large amount of raw data, which need to be processed and analyzed before they can be converted into useful environmental parameters. For example, it is necessary to remove noise interference and correct sensor errors through algorithms. After data processing, the system will finally output the specific environmental parameters of each environmental area.
[0077] Specifically, the system collected a large amount of raw data for each area. By removing noise interference and correcting sensor errors, the system obtained more accurate noise levels and echo intensity values. After data processing, the system output the specific environmental parameters for three areas in the indoor space: the living room, the bedroom, and the kitchen. For example, the noise level in the living room is 45 dB(A), the echo intensity is -10 dB, the light intensity is 300 lx, and the temperature is 25 °C; the noise level in the bedroom is 35 dB(A), the echo intensity is -15 dB, the light intensity is 150 lx, and the temperature is 23 °C; the noise level in the kitchen is 55 dB(A), the echo intensity is -5 dB, the light intensity is 500 lx, and the temperature is 27 °C. Based on these environmental parameters, the system can further analyze the interaction effects of each area and adjust the device configuration or interaction strategy as needed. For example, if it is found that the noise level in the kitchen is too high, the system will recommend that the user reduce the background noise or adjust the device position when using the intelligent voice assistant.
[0078] Therefore, multiple interactions are carried out on multiple environmental parameters, the spatial positions of multiple environmental areas, and the location of the voice device; a first interaction influence parameter is determined based on multiple environmental parameters and the spatial positions of multiple environmental areas, and a second interaction influence parameter is determined based on multiple environmental parameters and the location of the voice device. The first interaction influence parameter and the second interaction influence parameter are introduced to facilitate further management and control of the interaction space.
[0079] At this time, multiple environmental parameters (such as noise level, echo intensity, light intensity, temperature, etc.), the spatial positions of multiple environmental areas, and the location of the voice device are comprehensively considered for multiple interaction analysis. This means that the system needs to evaluate the impact of different environmental parameters on the interaction effects in different areas, as well as the interaction between these parameters and the location of the voice device.
[0080] First, the multiple environmental parameters, the spatial position information of the environmental areas, and the location information of the voice device collected need to be integrated to form a comprehensive data set. Interaction analysis is performed on these data, which includes evaluating the direct impact of different environmental parameters on the interaction effects (such as the impact of the noise level on the speech recognition accuracy), as well as the interaction between these parameters (such as the impact of light intensity and temperature on people's emotions, and then on the interaction experience). In the interaction analysis, the system also needs to consider the spatial position of the environmental area and the location of the voice device. For example, if the voice device is located in an area with a strong echo, this will have a negative impact on speech recognition. Similarly, if the device is located in an area with insufficient light, it will be difficult for the user to see the display interface of the device, thus affecting the interaction experience.
[0081] Based on the above multiple interaction analysis, two key interaction influence parameters are determined: the first interaction influence parameter and the second interaction influence parameter.
[0082] The first interaction impact parameter:
[0083] This is comprehensively determined based on multiple environmental parameters and the spatial positions of multiple environmental regions. It reflects the overall interaction effect of the environmental regions, taking into account the interactions between different environmental parameters and their direct impacts on the interaction. For example, an area with high noise, strong echoes, and insufficient light has a lower value of the first interaction impact parameter.
[0084] The second interaction impact parameter:
[0085] This is determined based on multiple environmental parameters and the location of the voice device. It reflects the interaction effect of the voice device in the current environment, especially considering the sensitivity of the device location to environmental parameters. For example, if the voice device is located in an area with a low noise level but strong echoes, the second interaction impact parameter will be negatively affected by the echoes and decrease.
[0086] Specifically, assume that in a large conference room, there is an intelligent voice control system for meeting recording and discussion. The conference room is divided into three areas: Area A (near the front door, with strong light but high noise level), Area B (in the center of the conference room, with moderate light and noise levels), and Area C (near the window, with strong light but affected by external noise). First, the environmental parameters (such as noise level, echo intensity, light intensity, etc.) of each area in the conference room and the location of the intelligent voice control system (assumed to be in Area B) are collected. Further multiple interaction analyses are carried out, and it is found that in Area A, due to the high noise level, it has a negative impact on speech recognition; in Area C, although the light is strong, it is affected by external noise. While in Area B, it has relatively suitable environmental conditions. Based on the above analysis, the system determines the first interaction impact parameter: lower in Area A, moderate in Area B, and slightly lower in Area C due to being affected by external noise. At the same time, since the intelligent voice control system is located in Area B, the system also determines the second interaction impact parameter: it has a better interaction effect in Area B, but is reduced due to environmental conditions in Area A and Area C.
[0087] Furthermore, based on the first interaction impact parameter, the second interaction impact parameter, and the spatial size of the interaction space, the interaction impact level of the interaction space is determined.
[0088] At this time, the first interaction influence parameter (reflecting the overall interaction effect of the environmental area) and the second interaction influence parameter (reflecting the interaction effect of the voice device in the current environment) are comprehensively evaluated. This usually involves weighted summation of the two parameters or using other appropriate comprehensive evaluation methods to obtain a single index reflecting the overall interaction effect of the interaction space. After obtaining the comprehensive evaluation index, the space size of the interaction space also needs to be considered. The space size is one of the important factors affecting the interaction effect. A larger space means that sound travels farther and the echo is stronger, while a smaller space is more easily affected by noise interference. Therefore, it is necessary to adjust the comprehensive evaluation index according to the space size to more accurately reflect the interaction influence level of the interaction space. Finally, based on the comprehensive evaluation index and the adjusted space size factor, the interaction influence level of the interaction space can be determined. This usually involves comparing the comprehensive evaluation index with a preset level threshold to determine which level the interaction space belongs to. The levels can be set to different levels such as high, medium, and low, so that users or system administrators can take corresponding optimization measures according to the interaction influence level.
[0089] Optionally, construct an interaction influence level learning model, taking the first interaction influence parameter, the second interaction influence parameter, and the space size of the interaction space as input features, and the interaction influence level as the output target. At this time, collect the first interaction influence parameter, the second interaction influence parameter, the space size, and the corresponding interaction influence level data of different interaction spaces, and preprocess the data, such as normalization, standardization, etc., to improve the effect of model training. Select the first interaction influence parameter, the second interaction influence parameter, and the space size of the interaction space as input features, and use the preprocessed data to train the interaction influence level learning model. Use the trained interaction influence level learning model to predict new data to obtain the interaction influence level.
[0090] In addition, construct an interaction influence level matching table to match different first interaction influence parameters, second interaction influence parameters, and the space size range of the interaction space with the corresponding interaction influence levels. At this time, define the threshold ranges of the first interaction influence parameter, the second interaction influence parameter, and the space size of the interaction space, assign a corresponding interaction influence level to each threshold range, and organize the defined matching rules into a table form, that is, the interaction influence level matching table. Each row in the interaction influence level matching table represents a threshold range and its corresponding interaction influence level. When it is necessary to determine the interaction influence level of a certain interaction space, query its first interaction influence parameter, second interaction influence parameter, and space size, find the threshold range matching these parameters in the interaction influence level matching table, and determine the corresponding interaction influence level. The interaction influence level matching table is as follows:
[0091] First interaction influence parameter range Second interaction influence parameter range Spatial size range Interaction influence level High High Small A (excellent) Medium Medium Medium B (good) Low Low Large C (poor) ... ... ... ...
[0092] In step S13, determine the interaction path of the voice device based on the location of the voice device and the location of the user.
[0093] In the specific implementation process of the present invention, the specific steps may be as follows:
[0094] S131: Obtain the location of the voice device and the interaction space of the voice device.
[0095] S132: Obtain the spatial image of the interaction space of the real-time voice device, and determine the spatial location of the user according to the recognition of the spatial image of the interaction space.
[0096] S133: Determine the location of the user based on the location of the voice device, the spatial system of the interaction space of the voice device, and the spatial location of the user.
[0097] S134: Determine the interaction path of the voice device according to the location of the voice device, the location of the user, and the obstacles in the interaction space.
[0098] In the embodiment of the present application, obtaining the location of the voice device and the interaction space of the voice device; obtaining the spatial image of the interaction space of the real-time voice device, and determining the spatial location of the user according to the recognition of the spatial image of the interaction space, ensures the accuracy of the spatial location of the user, so as to facilitate the subsequent management and control of the spatial location of the user.
[0099] First, introduce the location of the voice device and the interaction space of the voice device. In the interaction space of the voice device, obtain the spatial image of the interaction space in real time, which can be achieved through devices such as cameras, depth cameras, and infrared sensors installed in the interaction space. The obtained spatial image data needs to be processed and analyzed to identify the location of the user in the interaction space.
[0100] For the spatial image of the interaction space of the voice device, identify the spatial image of the interaction space of the voice device. At this time, extract image features for each position in the interaction space. These features can be local feature descriptors such as SIFT and SURF. Store the extracted features in a location matching table. Each feature corresponds to a specific spatial location. In the real-time recognition process, use the same feature extraction algorithm to extract features from the images captured by the camera, compare the extracted real-time features with the features in the location matching table, find the most similar features and their corresponding spatial locations, and determine the location information of the user in the current interaction space according to the matching result.
[0101] In addition, a pre-set location learning model is collected. By collecting a large amount of image data of users in the interaction space and annotating the location information of users in each image, the location of the user in the interaction space can be predicted based on the pre-set location learning model. In practical applications, when a user interacts with a voice device in the interaction space, the camera will capture the image of the user and input the image into the pre-set location learning model. The pre-set location learning model can output the location information of the user, so as to achieve precise control and interaction of the voice device.
[0102] Furthermore, the location of the user is determined based on the location of the voice device, the spatial system of the interaction space of the voice device, and the spatial location of the user, taking into account the overall situation of the location of the voice device, the spatial system of the interaction space of the voice device, and the spatial location of the user, ensuring accurate control of the location of the user.
[0103] At this time, the location of the voice device, the spatial system of the interaction space of the voice device, and the spatial location of the user are introduced. For the spatial system of the interaction space of the voice device, the spatial system of the interaction space includes information such as the shape, size, layout, and obstacle location of the interaction space.
[0104] Optionally, collect a large amount of training data containing the location of the voice device, the layout of the interaction space, and the location information of the user. The data should cover different types of interaction spaces (such as homes, offices, shopping malls, etc.) and user activity patterns. Extract the features of the voice device location, the spatial system of the interaction space (such as the location of walls, furniture, etc.), and the user location from the training data. Use a deep learning framework (such as TensorFlow, PyTorch) to build a location fusion model. The location fusion model should be able to receive the features of the voice device location, the layout of the interaction space, and the user location as inputs and output the location of the user.
[0105] In practical applications, input the features of the voice device location, the layout of the interaction space, and the user location obtained in real time into the trained location fusion model. The location fusion model outputs the location of the user. Optionally, in an intelligent home scenario, use an intelligent robot as the voice device. The home layout is known, and a camera is installed in the home to capture the user's location. By collecting a large amount of data on family members interacting with the intelligent robot in different rooms and at different locations, and extracting the features of the voice device location, the home layout, and the user location, a location fusion model can be trained to predict the precise location of the user in the home. In practical applications, when a family member interacts with the intelligent robot, the location fusion model can output the location of the user based on the data obtained in real time.
[0106] Therefore, by determining the interaction path of the voice device based on the location of the voice device, the location of the user, and the obstacles in the interaction space, multi-dimensional interaction among the location of the voice device, the location of the user, and the obstacles in the interaction space is achieved, ensuring the accuracy of the interaction path of the voice device.
[0107] At this time, the location of the voice device is the starting point of the interaction path. Accurately knowing the location of the voice device helps to plan the direction and intensity of sound propagation, and the location of the user is the end point of the interaction path. By tracking the user's location in real time, the system can dynamically adjust the sound propagation path to ensure that information can be directly and clearly transmitted to the user, while the obstacles in the interaction space will affect the sound propagation. It is necessary to identify and consider the location and shape of these obstacles in order to plan the best path to bypass or penetrate the obstacles.
[0108] In the multi-dimensional interaction among the location of the voice device, the location of the user, and the obstacles in the interaction space, the interaction between the location of the voice device and the location of the user, the interaction between the location of the voice device and the obstacles in the interaction space, and the interaction between the location of the user and the obstacles in the interaction space are introduced.
[0109] Regarding the interaction between the location of the voice device and the location of the user, the relative position relationship between the user and the voice device is mainly concerned. Knowing the user's current location and the location of the voice device to evaluate key information such as the distance and direction between the user and the device, which can usually be achieved through the positioning function in the voice device (such as the positioning service of the mobile phone APP, Bluetooth beacon, etc.).
[0110] Regarding the interaction between the location of the voice device and the obstacles in the interaction space, the position relationship between the voice device and the obstacles is concerned. The system needs to identify the obstacles in the interaction space (such as furniture, walls, etc.) and determine their impact on the sound propagation of the voice device.
[0111] Regarding the interaction between the location of the user and the obstacles in the interaction space, the position relationship between the user and the obstacles is concerned. The system needs to know the user's current location and the surrounding obstacle situation to evaluate the obstacles encountered when the user moves, which helps the system to provide accurate path guidance for the user and avoid the user colliding with the obstacles or getting into an impassable area.
[0112] After comprehensively considering the interaction relationships among the above three, the system can start to plan the interaction path of the voice device, and this process usually involves the following steps:
[0113] Data collection: The system needs to collect data on the user's location, the location of the voice device, and the location of obstacles in the interaction space. This data can be obtained through devices such as sensors, cameras, and positioning services in the smart home system.
[0114] Environmental modeling: Based on the collected data, the system can construct a three-dimensional model of the interaction space, including the location information of the user, the voice device, and the obstacles. This helps the system to more intuitively understand the spatial layout and obstacle distribution.
[0115] Path planning: Based on the three-dimensional model of the interaction space, the system can use path planning algorithms (such as the Dijkstra algorithm) to plan the optimal path from the user to the voice device (or from the voice device to the user). The algorithm should consider factors such as the user's movement speed, the location of the obstacles, and the characteristics of sound propagation.
[0116] Path optimization: The initially planned path needs to be optimized to improve the smoothness of the path, reduce the path length, or avoid potential interference sources. This can be achieved by adjusting algorithm parameters, adding constraint conditions, or using heuristic search and other methods.
[0117] In step S14, determine the virtual interaction mechanism of the voice device according to the interaction path of the voice device, the interaction influence level of the interaction space, and the service life of the voice device;
[0118] In the specific implementation process of the present invention, the specific steps may be:
[0119] S141: Obtain the interaction path of the voice device;
[0120] S142: Determine the interaction area based on the matching of the interaction space of the voice device and the interaction path of the voice device;
[0121] S143: Determine the service life of the voice device based on the age detection of the voice device;
[0122] S144: Perform multiple interactions on the interaction area, the interaction influence level of the interaction space, and the service life of the voice device;
[0123] S145: Determine the virtual interaction mechanism of the voice device according to the multiple interactions of the interaction area, the interaction influence level of the interaction space, and the service life of the voice device.
[0124] In the embodiment of the present application, obtain the interaction path of the voice device; determine the interaction area based on the matching of the interaction space of the voice device and the interaction path of the voice device.
[0125] At this time, the interaction space and interaction path of the voice device are introduced. After determining the interaction path, the matching relationship between the interaction space and the interaction path is further analyzed to determine an effective interaction area. The interaction area is the spatial range where the user interacts with the voice device. It should be large enough to ensure that the user can clearly hear the output of the voice device within this area, and the voice device can also accurately receive the user's input. To determine the interaction area, the following factors need to be considered:
[0126] 1. Interaction path and spatial layout: The system needs to analyze the matching degree between the interaction path and the interaction space layout to ensure that each point on the path is within an effective interaction range.
[0127] 2. Sound propagation characteristics: The system also needs to consider the sound propagation characteristics in the interaction space, such as sound attenuation, reflection, etc. This helps to determine an optimal interaction distance and angle.
[0128] 3. Obstacle influence: The system needs to evaluate the influence of obstacles on the interaction area. For example, if obstacles (such as walls or large furniture) block the sound propagation path, then the interaction area needs to be adjusted accordingly.
[0129] Specifically, after determining the interaction path, the corridor layout from the bedroom to the living room and the sound propagation characteristics are further analyzed. Considering the width of the corridor and the placement position of the furniture, the system determines an effective interaction area. This area covers the entire path from the bedroom door to in front of the intelligent robot in the living room, and ensures that the user can clearly hear the output of the intelligent robot at each position within this area. At the same time, the system also considers the attenuation and reflection characteristics of sound in the corridor to ensure that the volume and sound quality of the intelligent robot meet the user's requirements.
[0130] Furthermore, based on the age detection of the voice device, the service life of the voice device is determined, and the service life of the voice device is introduced to facilitate further management and control of the service life of the voice device.
[0131] At this time, check the physical condition of the voice device, including the sensitivity, integrity of the microphone and speaker, and whether there is physical damage. Through the log of the voice device or user reports, collect the usage frequency of the device, including the number of activations per day, call duration, etc. At the same time, evaluate the environment where the voice device is located, such as temperature, humidity, dust level, etc. These factors will affect the life of the device
[0132] Furthermore, compare the collected device information with historical data (such as the usage of other devices of the same model) to identify potential trends or patterns, analyze the failure rates of devices of the same model within different usage years to determine the approximate range of device lifespan. At the same time, comprehensively evaluate the collected device information, the comparison results of historical data, and the evaluation results of environmental factors to determine the approximate usage years of the device and provide the evaluation results of the device's usage years to the user.
[0133] Therefore, perform multiple interactions on the interaction area, the interaction impact level of this interaction space, and the usage years of the voice device; determine the virtual interaction mechanism of the voice device based on the multiple interactions of the interaction area, the interaction impact level of this interaction space, and the usage years of the voice device, which takes into account the overall consideration of the interaction path of the voice device, the interaction impact level of this interaction space, and the usage years of the voice device, improves the accuracy of the virtual interaction mechanism of the voice device, and enables the voice device to perform smooth online interactions with users in various scenarios.
[0134] At this time, comprehensively consider these three factors: the interaction area, the interaction impact level of the interaction space, and the usage years of the voice device, and perform multiple interaction analyses. This includes analyzing their mutual relationships, mutual constraints, and synergistic effects. For example, the size and shape of the interaction area will affect the propagation path and effect of sound in the interaction space; the interaction impact level of the interaction space will limit the effective range of the interaction area; and the usage years of the voice device will determine the performance of the device under specific interaction conditions.
[0135] Specifically, collect a large amount of data on the interaction area, the interaction impact level of the interaction space, and the usage years of the voice device. These data can come from actual user feedback, laboratory tests, or public databases. The collected data needs to be preprocessed, including data cleaning, normalization processing, etc., to ensure the quality and consistency of the data. Based on the data preprocessing, features useful for determining the virtual interaction mechanism need to be extracted from the original data. These features include the shape and size of the interaction area, the background noise level, echo and reverberation characteristics of the interaction space, the performance parameters of the microphone and speaker of the voice device, and the usage years, etc.
[0136] To collect a pre-set machine learning model, pre-processed data and extracted features are required to train the model. During the training process, the parameters and structure of the model need to be continuously adjusted to improve the prediction accuracy and generalization ability of the model. Once the model training is completed, it can be used to predict the virtual interaction mechanism under different interaction areas, interaction spaces, and service life. Specifically, new interaction areas, interaction spaces, and service life can be used as inputs, and the model can predict the optimal interaction strategy, acoustic processing algorithm, and device performance monitoring and maintenance plan, etc. These prediction results will be used as the basis for determining the virtual interaction mechanism of the voice device.
[0137] For the virtual interaction mechanism of the voice device, when the user issues a voice command, the voice device will capture and analyze these commands.
[0138] Based on the user's command and the current environmental conditions (such as the size of the interaction area, the background noise level of the interaction space, etc.), the machine learning model will predict the optimal interaction strategy. For example, if the user says "Turn on the lights in the living room", the machine learning model will determine the most appropriate light brightness and color according to the size, location of the living room, and the current light conditions. At the same time, the machine learning model will also consider the service life of the device to ensure the best interaction experience within the allowable range of device performance.
[0139] In this example, the virtual interaction mechanism of the voice device realizes the following functions:
[0140] Accurate recognition: Through the training of speech recognition technology and the machine learning model, the voice device can accurately recognize the user's voice commands.
[0141] Intelligent prediction: Based on the extracted features and the results predicted by the model, the voice device can intelligently predict the user's intentions and the device's response strategies.
[0142] Dynamic adjustment: According to the current environmental conditions and device status, the voice device can dynamically adjust the interaction strategy to ensure the best interaction experience.
[0143] Continuous optimization: Through user feedback and continuous data collection, the virtual interaction mechanism of the voice device can be continuously optimized and improved to adapt to different usage scenarios and user needs.
[0144] In addition, in the scenario of educational consulting interaction, further control over sensitive words is introduced. The voice commands issued by the user are introduced, and the corresponding role characteristics of the user are matched. Considering sensitive words, the voice commands issued by the user, and role characteristics comprehensively, more intelligent and personalized voice interaction is further realized, enabling the virtual assistant to provide appropriate feedback according to different roles and scenarios, while avoiding potential problems caused by sensitive words, and outputting the processed voice information for feedback to the user. By introducing the preset role description of the human voice and the sensitive word processing mechanism, a personalized interaction experience can be provided according to the needs and scenarios of different users, while avoiding the leakage or inappropriate expression of sensitive information, and enhancing the security and compliance of the interaction.
[0145] Specifically, first, the system receives the input voice data stream from the user. The system identifies the user's role (such as "master", "visitor", etc.) according to the preset role description, and adjusts the voice processing strategy according to the role characteristics. The system detects sensitive words in the input voice and replaces or transforms them according to the preset rules to ensure the compliance and security of the voice content. After role recognition and sensitive word processing, the system generates an output voice data stream for feedback to the user.
[0146] In addition, the system receives the input voice data stream. According to the required voice format (such as WAV, MP3, etc.), the system selects the corresponding encoder. According to the required voice format (such as WAV, MP3, etc.), the system selects the corresponding encoder. The system quantizes the voice data according to the bit depth to match the accuracy of the target format. After formatting, the system generates a voice file for storage or transmission. Optionally, voice encoding uses the Fourier transform (such as FFT) to analyze and compress the voice signal. Resampling uses interpolation algorithms, such as linear interpolation or spline interpolation, to change the sampling rate of the voice data. Quantization uses uniform quantization or non-uniform quantization methods to discretize the voice signal according to the bit depth. Further, the dialogue voice format processing function can convert the input voice into multiple standard formats and optimize it according to the sampling rate and bit depth to ensure the compatibility and operability of the voice data between different devices and systems, facilitating subsequent storage, transmission, and playback operations.
[0147] In step S15, in the virtual interaction mechanism of the voice device, a virtual interaction data set is determined based on the video image and voice data collected by the voice device; the virtual interaction content of the user and the virtual interaction content of the environment are output according to the recognition of the virtual interaction data set;
[0148] In the specific implementation process of the present invention, the specific steps may be:
[0149] S151: Obtain the virtual interaction mechanism of the voice device and trigger the interaction of the voice device with the user based on this virtual interaction mechanism;
[0150] S152: In the virtual interaction mechanism of the voice device, determine the video picture collected by the voice device based on the positioning and shooting of the voice device and the user.
[0151] S153: In the virtual interaction mechanism of the voice device, determine the voice data collected by the voice device based on the positioning and collection of the voice device and the user.
[0152] S154: Determine the virtual interaction data set according to the video picture and voice data collected by the voice device.
[0153] S155: Traverse the virtual interaction data set, and mark the data types of each virtual interaction data during the traversal.
[0154] S156: Form corresponding virtual interaction data combinations according to the classification of the data types of each virtual interaction data.
[0155] S157: Output the virtual interaction content of the user and the virtual interaction content of the environment according to the recognition of each virtual interaction data combination.
[0156] In the embodiments of the present application, obtain the virtual interaction mechanism of the voice device, and trigger the interaction of the voice device with the user based on this virtual interaction mechanism, realizing the interaction of the voice device with the user and ensuring the interaction effect of the voice device with the user.
[0157] At this time, obtain the virtual interaction mechanism of the voice device from development documents, configuration files, software updates, etc. The virtual interaction mechanism of the voice device refers to the way of information exchange and response between the voice device and the user through preset rules and algorithms. For the trigger of this virtual interaction mechanism, the user issues a voice command, the voice device captures the voice command and transmits it to the voice recognition module. The voice recognition module converts the voice signal into text information, and the natural language processing module understands and analyzes the text information. The instruction execution module performs corresponding operations according to the results of the understanding and analysis.
[0158] Specific interaction process:
[0159] The user walks into the living room and says to the intelligent robot: "Hi, Xiaozhi, play a song by singer A."
[0160] The intelligent robot captures the user's voice command and transmits it to the voice recognition module.
[0161] The voice recognition module converts the voice signal into text information: "Play a song by singer A."
[0162] The natural language processing module understands that the user's intention is to play a song by singer A.
[0163] The instruction execution module calls the music playback service and starts playing the songs of singer A.
[0164] Meanwhile, for further control of voice noise, the system receives the original voice data stream. The system identifies and removes the noise exceeding the background noise threshold according to the background noise threshold. The system divides the voice signal into independent voice segments according to the minimum recognition cutting duration to avoid misidentifying short-time noise as voice. After noise reduction and cutting processing, the system generates the processed voice data stream. Optionally, a filter such as a low-pass filter or a band-pass filter is used for noise reduction to remove high-frequency or specific-band noise. Energy detection or zero-crossing rate and other methods are used for voice segmentation to identify the boundaries of voice segments.
[0165] As follows: The functional relationship of the voice noise processing function:
[0166] Taudio = AudFilf(Saudio, Thnoise, Mtime) where:
[0167] Taudio: The processed voice data stream, that is, the voice signal after noise reduction processing.
[0168] Saudio: The original voice data stream, that is, the unprocessed user input voice.
[0169] Thnoise: The background noise threshold, used to set the allowed background noise level, and the noise exceeding this threshold will be filtered out.
[0170] Mtime: The minimum recognition cutting duration, used to determine the minimum duration of the effective voice segment in the voice signal to avoid misidentifying short-time noise as voice.
[0171] In addition, the system receives the original voice data stream. The system identifies the silent time between voice segments according to the dialogue intermittent tone determination duration to divide different dialogue segments. The system ignores the voice signals below the sound loudness filtering range to remove background noise or invalid voice. After segmentation processing, the system generates the segmented voice data set for subsequent segment-by-segment processing and analysis.
[0172] As follows: The functional relationship of the voice segmentation processing function:
[0173] Caudio = VoCf(Saudio, Dsli, Vrange) where:
[0174] Caudio: The segmented voice data set, that is, the set of voice segments after dividing the original voice data according to the dialogue interval and sound loudness.
[0175] Saudio: Original voice data stream.
[0176] Dsli: Dialogue interval determination duration, used to determine the silence time between two voice segments. If the duration exceeds this value, they are considered to be different dialogue segments.
[0177] Vrange: Sound loudness filter range, used to set the effective loudness range of voice signals. Voice signals below this range will be ignored.
[0178] Through the speech noise processing function and the speech segmentation processing function, the background noise can be effectively removed and the speech signal can be segmented, so that the system can more accurately recognize the user's voice commands, reduce misrecognition and missed recognition, and thus improve the overall accuracy of voice interaction.
[0179] The voice segmentation processing function divides the voice signal into multiple independent voice segments through reasonable segmentation, which makes it easier for the system to process and analyze each segment, thereby improving the efficiency of voice interaction, reducing system processing delays, and improving user experience.
[0180] Furthermore, in the virtual interaction mechanism of the voice device, the video images captured by the voice device are determined based on the positioning and shooting of the voice device and the user.
[0181] At this time, the virtual interaction mechanism of the voice device is controlled in real time, and the voice device and the user's positioning shooting are introduced. The voice device can automatically adjust the camera's framing according to the user's actual location and voice interaction needs, thereby providing a more intuitive and convenient interactive experience.
[0182] Optionally, after determining the user's position, the voice device (if a camera is integrated) automatically adjusts the direction, angle and focal length of the camera to ensure that the user is in a suitable position in the picture. The camera starts to capture the video according to the adjusted parameters. The collected video data is encoded and processed and is ready for transmission or storage. At the same time, the processed video can be directly displayed to the user through the display of the voice device (if the device has a display). Alternatively, the video can be transmitted to other devices (such as smartphones, tablets or TVs) through a network connection for display.
[0183] Furthermore, in the virtual interaction mechanism of the voice device, the voice data collected by the voice device is determined based on the positioning collection of the voice device and the user.
[0184] At this time, the virtual interaction mechanism of the voice device is controlled in real time, and the positioning shooting of the voice device and the user is introduced. The voice device can provide a more intuitive and convenient interactive experience based on the user's actual location and voice interaction needs.
[0185] Optionally, once the user's location is determined, the voice device adjusts its acquisition parameters according to a preset strategy, such as the sensitivity, sampling rate, directivity, etc. of the microphone, to ensure that the user's voice is captured from the best angle and distance. According to the formulated strategy, the voice device starts to collect the user's voice data.
[0186] During the acquisition process, the voice device uses technologies such as automatic gain control (AGC), noise suppression (NS), and echo cancellation (AEC) to optimize the voice quality; after the acquired voice data is preprocessed (such as filtering, denoising), it is ready for subsequent speech recognition or transmission. The processed voice data can be played through the built-in speaker of the voice device (for real-time feedback), or transmitted through a network connection to other devices (such as smartphones, servers) for further processing or storage.
[0187] Meanwhile, a virtual interaction data set is determined based on the video images and voice data collected by the voice device; the virtual interaction data set is traversed, and the data types of each virtual interaction data are marked during the traversal process to facilitate subsequent control of the data types of each virtual interaction data.
[0188] At this time, the video images and voice data collected by the voice device are introduced, and the video images and voice data collected by the voice device are aggregated to output a virtual interaction data set. Optionally, the voice device first integrates the video image data collected from the camera and the voice data collected from the microphone. These data are combined into a comprehensive data stream or data packet, forming the initial form of the virtual interaction data set. The system performs preliminary processing on the integrated data to identify and extract key interaction information, including the user's facial features, gesture movements, voice content, voice rhythm, etc. These together constitute the virtual interaction data set.
[0189] Furthermore, start traversing this data set, checking and analyzing each data item. The traversal process ensures that each element in the virtual interaction data set is fully considered and processed. At the same time, during the traversal process, the system classifies and marks each virtual interaction data. For example, video image data is marked as "visual data", and voice data is marked as "audio data". More detailed marks include "facial expression data", "gesture data", "voice command data", etc. These marks help the system to process and analyze the data more effectively subsequently.
[0190] Specifically, the intelligent robot integrates the user video images collected from the camera and the user voice collected from the microphone, identifies and extracts key interaction information such as the user's facial expressions, gesture actions, and voice content, forms a virtual interaction data set, and starts traversing this data set to check and analyze each data item (such as facial expressions, gesture actions, voice content). During the traversal process, the system classifies and tags each virtual interaction data. For example, facial expression data is tagged as "facial data", gesture action data is tagged as "gesture data", and voice content data is tagged as "voice command data".
[0191] Therefore, corresponding virtual interaction data combinations are formed according to the classification of the data types of each virtual interaction data; the virtual interaction content of the user and the virtual interaction content of the environment are output according to the recognition of each virtual interaction data combination, and the virtual interaction content of the user and the virtual interaction content of the environment are introduced to facilitate the refined control of the virtual interaction content of the user and the virtual interaction content of the environment.
[0192] At this time, the data types of each virtual interaction data (such as visual data, audio data, gesture data, etc.) are introduced, classified according to the data types of each virtual interaction data (such as visual data, audio data, gesture data, etc.), and the classified data is combined into different virtual interaction data combinations. Each combination contains specific types of data, and these data together reflect the interaction information of a certain aspect of the user or the environment. Each virtual interaction data combination is recognized and analyzed to extract the virtual interaction content of the user or the environment. For example, the user's facial expressions, body postures, etc. are recognized from the visual data combination; the user's voice commands, emotional intonations, etc. are recognized from the audio data combination; the user's gesture actions, pointing, etc. are recognized from the gesture data combination.
[0193] Furthermore, the identified user virtual interaction content and environment virtual interaction content are output for subsequent processing or response. The output forms include various formats such as text, image, and audio, specifically depending on the system requirements and implementation methods. At the same time, after introducing the user's virtual interaction content and the environment virtual interaction content, the system can perform refined control on these contents. For example, the system can adjust the interaction method according to the user's facial expressions and emotional intonations to provide a more user-friendly service; or optimize the accuracy of speech recognition according to environmental conditions such as light and sound.
[0194] For the recognition of each virtual interaction data combination, after the voice device collects the user's voice data, video picture data, and other sensor data, these data will be classified according to their types and characteristics to form different virtual interaction data combinations, including the user's voice command combination, facial expression combination, gesture action combination, environmental sound combination, light condition combination, etc.
[0195] Optionally, for the voice command combination, the system will use speech recognition technology for recognition to convert the user's voice into a text command; for the facial expression combination and gesture action combination, the system will use computer vision technology and machine learning algorithms for recognition and analysis to extract information such as the user's emotional state and gesture intention. At the same time, by recognizing the user's voice command combination, the system can output the content of the user's voice command, such as "turn on the light", "play music", etc. By recognizing the user's facial expression combination and gesture action combination, the system can output virtual interaction content such as the user's emotional state (such as happy, angry) and gesture intention (such as pointing, waving).
[0196] In step S16, determine the virtual interaction list according to the user's virtual interaction content, the interaction part of the voice device, the priority matching table of interaction actions, and the user's work schedule. At the same time, determine the expression factor of each interaction item in the virtual interaction list based on the virtual interaction content of the environment and the interaction part of the voice device;
[0197] In the specific implementation process of the present invention, the specific steps can be:
[0198] S161: Obtain the user's virtual interaction content;
[0199] S162: Determine multiple action modules based on the online detection of the voice device, and match the interaction part of the voice device according to the multiple action modules and the interaction matching table;
[0200] S163: Determine multiple action nodes of the user based on the recognition of the video picture collected by the voice device, and determine the user's real-time action according to the relative positions of the multiple action nodes and the action recognition module;
[0201] S164: Determine multiple interaction commands according to the multiple interactions of the user's virtual interaction content, the interaction part of the voice device, and the user's real-time action;
[0202] S165: Determine the virtual interaction list based on the interaction actions corresponding to the multiple interaction commands, the priority matching table of interaction actions, and the user's work schedule;
[0203] S166: Perform multiple interactions on the virtual interaction content of the environment and the interaction part of the voice device;
[0204] S167: Determine the expression factors of each interaction item in the virtual interaction list based on the virtual interaction content based on the environment and the multiple interactions of the interaction part of the voice device. At this time, define the expression coefficients of the corresponding interaction items based on each expression factor, and determine the expression actions of the voice device according to the expression coefficients, the expression matching table, and the expression factors.
[0205] In the specific implementation process of the present invention, obtain the user's virtual interaction content; determine multiple action modules based on the online detection of the voice device, and match the interaction part of the voice device according to the multiple action modules and the interaction matching table, realizing the dynamic matching of the multiple action modules and the interaction matching table, and ensuring the accuracy of the interaction part of the voice device.
[0206] At this time, the user's virtual interaction content is introduced. At the same time, the voice device detects its own functional state in real time, such as whether the microphone is working properly, whether the speech recognition engine is online, etc., and detects the user's current interaction environment, such as noise level, background sound, etc., to ensure the accuracy and reliability of the interaction. According to the online detection results, determine the available action modules. For example, if the microphone is working properly and the speech recognition engine is online, then determine that the speech recognition module is one of the available action modules.
[0207] At the same time, preset an interaction matching table, which records the matching relationships between different interaction contents and the corresponding action modules. For example, for the voice command "turn on the lights in the living room", the interaction matching table indicates that the speech recognition module should be used for command parsing, and the smart home control module should be called to execute the corresponding operation. According to the interaction matching table, match the user's virtual interaction content with the available action modules to determine the specific interaction part of the voice device.
[0208] Specifically, the user's voice command "turn on the lights in the living room" has been obtained. First, confirm through online detection that both the microphone and the speech recognition engine are in normal working states. Then, the system checks the interaction matching table and finds that the command "turn on the lights in the living room" is associated with the speech recognition module and the smart home control module. Therefore, the system determines to use the speech recognition module to parse the voice command and calls the smart home control module to execute the operation of "turn on the lights in the living room".
[0209] Furthermore, determine multiple action nodes of the user based on the recognition of the video images collected by the voice device, and determine the user's real-time actions according to the relative positions of the multiple action nodes and the action recognition module, taking into account the relative positions of the multiple action nodes and the action recognition module as a whole, and ensuring the accurate evaluation of the user's real-time actions.
[0210] At this time, multiple action nodes of the user are determined based on the recognition of the video images collected by the voice device, and the real-time action of the user is determined according to the relative positions of the multiple action nodes and the action recognition module; a virtual interaction list is determined based on the interaction actions corresponding to multiple interaction instructions, the priority matching table of interaction actions, and the user's work schedule, introducing multiple interactions of the interaction actions corresponding to multiple interaction instructions, the priority matching table of interaction actions, and the user's work schedule, and achieving precise control of the virtual interaction list.
[0211] Specifically, the voice device (such as a smart robot, a smart camera, etc.) is built with a camera that can collect the video images of the user's environment in real time. Using computer vision technology, the collected video images are analyzed and processed to identify multiple action nodes of the user. These action nodes include key parts such as the head, hands, and body. At this time, through image recognition algorithms, such as deep learning models (convolutional neural network CNN, etc.), feature extraction and classification are performed on the pixel information in the video images, so as to determine the action nodes of the user.
[0212] After identifying multiple action nodes of the user, the system further analyzes the relative position relationship between these action nodes. For example, the position of the hand relative to the position of the head can reflect whether the user is waving, pointing in a certain direction, or making other gestures. Using a preset action recognition module (including functions such as gesture recognition and facial expression recognition), the real-time action of the user is recognized and classified. These modules are usually trained based on deep learning algorithms and can accurately recognize various actions of the user. When determining the real-time action of the user, the system not only considers the information of a single action node, but also comprehensively considers the relative positions of multiple action nodes and the overall action pattern. This helps to improve the accuracy and robustness of action recognition.
[0213] Optionally, taking a smart robot as an example. After identifying the head and hand action nodes of the user, the system further analyzes the relative position relationship between these two action nodes. When the user makes a gesture of waving goodbye, the system can accurately recognize this action and, combined with the gesture recognition function in the action recognition module, classify it as a specific action of "waving goodbye". At the same time, the system also combines the facial expression recognition function to analyze the user's emotional state (such as smiling, the expression when waving goodbye, etc.), so as to more comprehensively understand the user's interaction intention.
[0214] Furthermore, multiple interaction instructions are determined based on the multiple interactions of the user's virtual interaction content, the interaction part of the voice device, and the user's real-time action; a virtual interaction list is determined based on the interaction actions corresponding to multiple interaction instructions, the priority matching table of interaction actions, and the user's work schedule;
[0215] At this time, the virtual interaction content is the instructions or requests input by the user through voice, text, or other means; the interaction part of the voice device is the part of the voice device responsible for processing the user input and generating responses, and the real-time actions of the user are the user actions captured and recognized by the system through a camera or other sensors, such as gestures, body postures, etc. Therefore, the true intention of the user and the required interaction instructions are determined by comprehensively considering the above three aspects of information.
[0216] Finally, the interaction instructions can be further refined and defined so that the system can more accurately understand the user's intention and make corresponding responses. The following are some examples of interaction instructions:
[0217] Device on / off instructions: such as "Turn on the lights in the bedroom", "Turn off the air conditioner in the living room".
[0218] Device status adjustment instructions: such as "Adjust the lights in the bedroom to the brightest", "Adjust the temperature of the air conditioner in the living room to 25 degrees".
[0219] Media playback instructions: such as "Play a piece of relaxing music", "Play the latest news".
[0220] Information query instructions: such as "Query today's weather forecast", "Display my schedule".
[0221] User feedback instructions: such as "I am very satisfied with this result", "I need more information".
[0222] Each interaction instruction should contain sufficient context information so that the system can accurately understand the user's intention and make corresponding responses. In addition, the system should also have the ability to learn and adapt, and can continuously optimize the recognition and response strategies of interaction instructions according to the user's feedback and behavior habits.
[0223] In addition, for the virtual interaction list, the interaction actions corresponding to multiple interaction instructions, the priority matching table of interaction actions, and the user's work schedule table are introduced, and multiple interactions are performed on the interaction actions corresponding to multiple interaction instructions, the priority matching table of interaction actions, and the user's work schedule table.
[0224] For the multiple interactions of the interaction actions corresponding to multiple interaction instructions, the priority matching table of interaction actions, and the user's work schedule table, at this time, for the interaction actions corresponding to multiple interaction instructions, when the user interacts with the system through voice, text, or other means, multiple interaction instructions will be generated. Each instruction corresponds to one or more specific interaction actions, and these actions are the direct results of the system executing the user's requests.
[0225] For the priority matching table of interaction actions, the priority matching table is used to determine which instruction or action should be executed first, and the corresponding priority is marked for the interaction action. The considerations include the urgency of the instruction, the importance of the user, the status of the device, etc.
[0226] Urgency: Sort the instructions according to the urgency of the instructions. For example, an emergency alarm usually has the highest priority.
[0227] User importance: Set the priority according to the user's identity or privilege level. For example, VIP users enjoy a higher service priority.
[0228] Device status: Consider the current usage of the device. For example, avoid disturbing the user when performing a critical task.
[0229] In addition, the user's work schedule provides the user's current and future work plans and status, which is crucial for the system to understand the user's real needs and make appropriate responses, paying attention to the user's current status, future plans, and preference settings.
[0230] Optionally, first receive the interaction instructions input by the user in the form of voice, text, etc. Parse the received instructions to determine the corresponding interaction actions. According to the preset priority matching table, sort multiple instructions, and combine with the user's work schedule to evaluate whether the current instruction conflicts with the user's work status or preferences. According to the sorting and evaluation results, execute the corresponding interaction actions and feedback the execution results to the user, including successful execution, execution failure, or the situation that requires the user to further confirm.
[0231] Specifically, assume that the user is working at home, and his work schedule shows that he has an important video conference about to start. At this time, the user requests the system to "play a piece of relaxing music" through a voice command, and at the same time he has a gesture pointing to the audio device.
[0232] At this time, the system recognizes the user's voice command and gesture action, determines that the intention is to play music. The system checks the priority matching table and finds that there is no urgent instruction that needs to be processed first. The system checks the user's work schedule and finds that the user is about to participate in an important video conference. Considering that the user does not want the music to interfere with the conference, the system evaluates that the instruction to play music does not match the user's current work status. Based on the above evaluation, the system decides not to execute the instruction to play music temporarily, but to send a reminder to the user, informing him of the upcoming video conference, and asking if he really needs to play music. The system feeds back the execution decision to the user and makes corresponding adjustments according to the user's further instructions. Therefore, through such a multi-interaction processing process, the system can more accurately understand the user's intention and make appropriate responses according to the user's current status and work schedule, so as to provide a more intelligent and personalized user interaction experience.
[0233] Meanwhile, perform multiple interactions on the virtual interaction content of the environment and the interaction part of the voice device; determine the expression factors of each interaction item in the virtual interaction list based on the multiple interactions of the virtual interaction content of the environment and the interaction part of the voice device. At this time, define the expression coefficients of the corresponding interaction items based on each expression factor, and determine the expression actions of the voice device according to the expression coefficients, the expression matching table, and the expression factors. At this time, the expression factors are introduced into the virtual interaction list, realizing the phased interaction actions and corresponding expression control of the voice device.
[0234] At this time, performing multiple interactions on the virtual interaction content of the environment and the interaction part of the voice device can greatly enrich the interaction experience between the user and the system. By introducing expression factors, the system can not only respond to the user's instructions but also interact with the user in a more vivid and user-friendly way.
[0235] Multiple interactions of the virtual interaction content of the environment and the interaction part of the voice device are introduced. The virtual interaction content of the environment includes the physical environment where the user is located and the environments simulated through technologies such as virtual reality (VR) and augmented reality (AR). The environment contains various interactive elements such as lights, sound systems, furniture, etc., as well as virtual characters, scenes, etc. The interaction part of the voice device is mainly a system component responsible for processing user inputs (such as voice commands) and generating responses (such as performing actions, providing information, etc.). Therefore, in the interaction between these two parts, the system needs to comprehensively consider the context information of the environment (such as the current environmental state, the user's location, etc.) and the user's voice commands to determine the user's true intention and the required interaction actions.
[0236] Based on the multiple interactions, the system can further analyze the virtual interaction content of the environment and the interaction part of the voice device to determine the expression factors of each interaction item in the virtual interaction list. These expression factors include:
[0237] Emotional color: Such as emotional states like happiness, sadness, surprise, etc., which can be inferred from the user's intonation, the environmental atmosphere, etc.
[0238] Characteristics of the interaction object: For example, whether the interaction object is serious and formal (such as a work report) or relaxed and pleasant (such as a family gathering), which will affect the interaction method and expression actions adopted by the system.
[0239] User preferences: Users prefer the system to interact with them in a certain specific way, such as being more friendly or more professional.
[0240] Furthermore, based on the determined expression factors, the system can define an expression coefficient for each interaction item. This coefficient is a quantitative index used to represent the intensity and type of the expression actions that should be adopted when executing this interaction item. For example, a happy expression coefficient indicates that the system interacts with the user in a relaxed and pleasant tone and actions. Further, according to the expression coefficient, the expression matching table, and the expression factors, the expression actions of the voice device are determined.
[0241] At this time, the expression matching table is a predefined table used to map the expression coefficient to specific expression actions. For example, a specific expression coefficient corresponds to actions such as smiling, nodding, and waving. The system combines the expression coefficient, the expression matching table, and the expression factors to determine the specific expression actions that should be adopted when executing each interaction item. After determining the expression actions of each interaction item, the system can execute these actions in stages and perform expression control during the execution process. This means that the system will dynamically adjust the expression actions according to the current interaction stage and the user's feedback to ensure the smoothness and naturalness of the entire interaction process.
[0242] In another embodiment, an expression matching table is collected, and the expression matching table is as follows:
[0243]
[0244]
[0245] For the expression coefficient, according to factors such as the user's voice intonation and the environmental atmosphere, the expression coefficient of the current interaction item is automatically determined. For example, if the user shows a happy mood and the environmental atmosphere is relaxed, the expression coefficient is determined to be 0.9. According to the determined expression coefficient, the corresponding expression action range is found in the expression matching table, and a specific expression action is selected. For example, when the expression coefficient is 0.9, "smiling" is selected as the expression action of the voice device among the example expression actions corresponding to "happy".
[0246] In addition, a preset expression learning model is collected. For a new interaction item, the system can automatically extract expression factors (such as the user's voice intonation and the environmental atmosphere) and input them into the trained expression learning model. The expression learning model will output the corresponding expression coefficient. According to the output expression coefficient, the system can search a predefined expression action library to generate specific expression actions. The expression action library contains templates or parametric representations of various expression actions, and the generation algorithm dynamically generates new expression actions according to the expression coefficient.
[0247] In an embodiment of the present invention, through the method in the embodiment of the present invention, an interaction space of a voice device is determined based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range; an interaction influence level of the interaction space is determined according to the environmental detection of the interaction space of the voice device; an interaction path of the voice device is determined based on the location of the voice device and the location of the user; a virtual interaction mechanism of the voice device is determined according to the interaction path of the voice device, the interaction influence level of the interaction space, and the service life of the voice device, which comprehensively considers the interaction path of the voice device, the interaction influence level of the interaction space, and the service life of the voice device, improves the accuracy of the virtual interaction mechanism of the voice device, and enables the voice device to perform smooth online interaction with the user in various scenarios.
[0248] Further, in the virtual interaction mechanism of the voice device, a virtual interaction data set is determined based on the video picture and voice data collected by the voice device; the virtual interaction content of the user and the virtual interaction content of the environment are output according to the recognition of the virtual interaction data set, and the virtual interaction content of the user and the virtual interaction content of the environment are introduced to facilitate the refined management and control of the virtual interaction content of the user and the virtual interaction content of the environment.
[0249] Therefore, a virtual interaction list is determined according to the virtual interaction content of the user, the interaction part of the voice device, the priority matching table of interaction actions, and the work schedule of the user. At the same time, the expression factors of each interaction item in the virtual interaction list are determined based on the virtual interaction content of the environment and the interaction part of the voice device, ensuring the accuracy of the virtual interaction list and introducing expression factors into the virtual interaction list, realizing the phased interaction actions of the voice device and the corresponding expression management and control.
[0250] Embodiment 4
[0251] In this embodiment, as Figure 3 shown, a virtual interaction device based on a voice device is provided, including:
[0252] An interaction space module 21, configured to determine an interaction space of a voice device based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range;
[0253] An interaction influence level module 22, configured to determine an interaction influence level of the interaction space according to the environmental detection of the interaction space of the voice device;
[0254] An interaction path module 23, configured to determine an interaction path of the voice device based on the location of the voice device and the location of the user;
[0255] The virtual interaction mechanism module 24 is used to determine the virtual interaction mechanism of the voice device according to the interaction path of the voice device, the interaction influence level of the interaction space, and the service life of the voice device;
[0256] The virtual interaction content module 25 is used to determine a virtual interaction data set based on the video picture and voice data collected by the voice device in the virtual interaction mechanism of the voice device; output the virtual interaction content of the user and the virtual interaction content of the environment according to the recognition of the virtual interaction data set;
[0257] The virtual interaction list module 26 is used to determine a virtual interaction list according to the virtual interaction content of the user, the interaction part of the voice device, the priority matching table of interaction actions, and the work schedule of the user. At the same time, determine the expression factors of each interaction item in the virtual interaction list based on the virtual interaction content of the environment and the interaction part of the voice device.
[0258] For the specific limitations of the virtual interaction device based on the voice device, reference can be made to the limitations of the virtual interaction method based on the voice device in the above text, which will not be elaborated here. Each unit in the above virtual interaction device based on the voice device can be implemented in whole or in part by software, hardware, and their combination. The above units can be embedded in or independent of the processor in the electronic device in the form of hardware, or stored in the memory of the electronic device in the form of software, so that the processor can call and execute the operations corresponding to the above units.
[0259] Embodiment Five
[0260] In this embodiment, an electronic device is provided. Its internal structure diagram can be as Figure 4 shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program, and a database is deployed on the non-volatile storage medium, and the database is used to store user behavior data and user portraits. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with other electronic devices on which application software is deployed. When the computer program is executed by the processor, it implements a virtual interaction method based on a voice device. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the electronic device, or an external keyboard, touchpad, or mouse, etc.
[0261] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0262] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A virtual interaction method based on a voice device, characterized in that, Including: Determine the interaction space of the voice device based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range; Determine the interaction influence level of the interaction space according to the environmental detection of the interaction space of the voice device; Determine the interaction path of the voice device based on the location of the voice device and the location of the user; Determine the virtual interaction mechanism of the voice device according to the interaction path of the voice device, the interaction influence level of the interaction space, and the service life of the voice device; In the virtual interaction mechanism of the voice device, determine the virtual interaction data set based on the video picture and voice data collected by the voice device; output the virtual interaction content of the user and the virtual interaction content of the environment according to the recognition of the virtual interaction data set; Determine the virtual interaction list according to the virtual interaction content of the user, the interaction part of the voice device, the priority matching table of interaction actions, and the work schedule of the user. At the same time, determine the expression factors of each interaction item in the virtual interaction list based on the virtual interaction content of the environment and the interaction part of the voice device.
2. The virtual interaction method based on a voice device according to claim 1, wherein The determination of the interaction space of the voice device based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range includes: In the indoor space, determine the voice device based on the traversal of the indoor space and mark the location of the voice device; Determine the voice collection module according to the online detection of the voice device, and determine the theoretical voice collection range of the voice device according to the traversal of the voice collection module; determine the actual voice collection range of the voice device according to the test of multiple voice collection nodes in the theoretical voice collection range of the voice device; Determine the video collection module according to the online detection of the voice device, and determine the theoretical video collection range of the voice device according to the traversal of the video collection module; determine the actual video collection range of the voice device according to the test of multiple video collection nodes in the theoretical video collection range of the voice device; Interact the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range, and output multiple spatial features according to the interaction of the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range; Determine the interaction space of the voice device according to multiple spatial features, the location of the voice device, and the location of the user.
3. The virtual interaction method based on a voice device according to claim 2, wherein The determination of the interaction influence level of the interaction space according to the environmental detection of the interaction space of the voice device includes: Obtain the interaction space of the voice device; Determine multiple spatial nodes based on the interaction space of the voice device and the multiple detections of the voice device; Determine multiple environmental areas according to multiple spatial nodes and the location of the voice device; Output multiple environmental parameters based on the environmental detection of multiple environmental areas; Perform multiple interactions on multiple environmental parameters, the spatial locations of multiple environmental areas, and the location of the voice device; Determine the first interaction influence parameter according to multiple environmental parameters and the spatial locations of multiple environmental areas, and determine the second interaction influence parameter according to multiple environmental parameters and the location of the voice device; Determine the interaction influence level of the interaction space based on the first interaction influence parameter, the second interaction influence parameter, and the spatial size of the interaction space.
4. The virtual interaction method based on a voice device according to claim 1, wherein, The determination of the interaction path of the voice device based on the location of the voice device and the location of the user includes: Obtain the location of the voice device and the interaction space of the voice device; Obtain the spatial image of the interaction space of the voice device in real time, and determine the spatial location of the user based on the recognition of the spatial image of the interaction space; Determine the location of the user based on the location of the voice device, the spatial system of the interaction space of the voice device, and the spatial location of the user; Determine the interaction path of the voice device according to the location of the voice device, the location of the user, and the obstacles in the interaction space.
5. The virtual interaction method based on a voice device according to any one of claims 1 to 4, characterized in that, The determination of the virtual interaction mechanism of the voice device according to the interaction path of the voice device, the interaction influence level of the interaction space, and the service life of the voice device includes: Obtain the interaction path of the voice device; Determine the interaction area based on the matching between the interaction space of the voice device and the interaction path of the voice device; Determine the service life of the voice device based on the age detection of the voice device; Perform multiple interactions on the interaction area, the interaction influence level of the interaction space, and the service life of the voice device; Determine the virtual interaction mechanism of the voice device according to the multiple interactions of the interaction area, the interaction influence level of the interaction space, and the service life of the voice device.
6. The virtual interaction method based on a voice device according to claim 1, characterized in that In the virtual interaction mechanism of the voice device, determine the virtual interaction data set based on the video frame and voice data collected by the voice device; Output the virtual interaction content of the user and the virtual interaction content of the environment according to the recognition of the virtual interaction data set, including: Obtain the virtual interaction mechanism of the voice device, and trigger the interaction of the voice device with the user based on the virtual interaction mechanism; In the virtual interaction mechanism of the voice device, determine the video frame collected by the voice device based on the positioning shooting of the voice device and the user; In the virtual interaction mechanism of the voice device, determine the voice data collected by the voice device based on the positioning collection of the voice device and the user; Determine the virtual interaction data set according to the video frame and voice data collected by the voice device; Traverse the virtual interaction data set, and mark the data types of each virtual interaction data during the traversal; Form corresponding virtual interaction data combinations according to the classification of the data types of each virtual interaction data; Output the virtual interaction content of the user and the virtual interaction content of the environment according to the recognition of each virtual interaction data combination.
7. The virtual interaction method based on a voice device according to claim 6, wherein The determination of the virtual interaction list according to the virtual interaction content of the user, the interaction part of the voice device, the priority matching table of interaction actions, and the work schedule of the user, and at the same time, determine the expression factors of each interaction item in the virtual interaction list based on the virtual interaction content of the environment and the interaction part of the voice device, includes: Obtain the virtual interaction content of the user; Determine multiple action modules based on the online detection of the voice device, and match the interaction part of the voice device according to the multiple action modules and the interaction matching table; Determine multiple action nodes of the user based on the recognition of the video images collected by the voice device, and determine the real-time action of the user according to the relative positions of the multiple action nodes and the action recognition module; Determine multiple interaction instructions according to the multiple interactions of the user's virtual interaction content, the interaction part of the voice device, and the user's real-time action.
8. The virtual interaction method based on a voice device according to claim 7, wherein Determine the virtual interaction list according to the user's virtual interaction content, the interaction part of the voice device, the priority matching table of interaction actions, and the user's work schedule. At the same time, determine the expression factors of each interaction item in the virtual interaction list based on the virtual interaction content of the environment and the interaction part of the voice device. It also includes: Determine the virtual interaction list based on the interaction actions corresponding to the multiple interaction instructions, the priority matching table of interaction actions, and the user's work schedule; Perform multiple interactions on the virtual interaction content of the environment and the interaction part of the voice device; Determine the expression factors of each interaction item in the virtual interaction list based on the multiple interactions of the virtual interaction content of the environment and the interaction part of the voice device. At this time, define the expression coefficient of the corresponding interaction item based on each expression factor, and determine the expression action of the voice device according to the expression coefficient, the expression matching table, and the expression factors.
9. A virtual interaction device based on a voice device, characterized in that, The virtual interaction device based on the voice device is applied to the virtual interaction method based on the voice device as described in any one of claims 1-8. The virtual interaction device based on the voice device includes: An interaction space module, configured to determine the interaction space of the voice device based on the location of the voice device, the actual voice collection range of the voice device, and the actual video collection range; An interaction influence level module, configured to determine the interaction influence level of the interaction space according to the environmental detection of the interaction space of the voice device; An interaction path module, configured to determine the interaction path of the voice device based on the location of the voice device and the location of the user; A virtual interaction mechanism module, configured to determine the virtual interaction mechanism of the voice device according to the interaction path of the voice device, the interaction influence level of the interaction space, and the service life of the voice device; A virtual interaction content module, configured to determine a virtual interaction data set based on the video images and voice data collected by the voice device in the virtual interaction mechanism of the voice device; output the user's virtual interaction content and the virtual interaction content of the environment according to the recognition of the virtual interaction data set; A virtual interaction list module, configured to determine the virtual interaction list according to the user's virtual interaction content, the interaction part of the voice device, the priority matching table of interaction actions, and the user's work schedule. At the same time, determine the expression factors of each interaction item in the virtual interaction list based on the virtual interaction content of the environment and the interaction part of the voice device.
10. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the virtual interaction method based on the voice device as described in any one of claims 1 to 8.