A Metaverse Intelligent Interaction Method and System

Through real-time acquisition and advanced algorithms, user interaction information is analyzed, voice and action interaction is recognized, and different types of interactions are carefully analyzed, which solves the interaction recognition problems caused by different user habits and improves user experience and recognition accuracy.

CN118689351BActive Publication Date: 2025-05-23BEIJING SURYEE SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410725439.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2025-05-23
Estimated Expiration
2044-06-05

AI Technical Summary

Technical Problem

Different users use intelligent interaction systems due to different habits, making it difficult for the system to recognize interactive information, reducing user interaction experience.

Method used

By collecting user interaction information in real time, using advanced recognition algorithms for processing and analysis, identifying voice interactions and action interactions, and analyzing voice interactions from both environment and language, and performing three-dimensional modeling and trajectory analysis for action interactions.

Benefits of technology

It improves the recognition accuracy during the interaction process, improves the user interaction experience, and ensures that the system can effectively identify and respond to user interaction information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118689351B_ABST
    Figure CN118689351B_ABST
Patent Text Reader

Abstract

The present invention discloses a metaverse intelligent interaction method and system thereof, which relates to the field of intelligent interaction technology, and solves the technical problem that different users may have problems in the interaction process due to different interaction modes when interacting, and further cannot interact well with users, thereby reducing the user interaction experience. The present invention identifies and analyzes the interaction information generated by the user, and classifies the interaction information, and separately analyzes and identifies voice interaction and action interaction. For voice interaction, the interaction content is identified and analyzed from the two aspects of environment and language to avoid interaction problems caused by environment and language. For action interaction, the user is analyzed by acquiring the video of the interaction process, and combined with multiple factors for analysis, the recognition accuracy in the interaction process is improved, and the user interaction experience is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent interaction technology, and specifically to an intelligent interaction method and system of a metaverse. Background Art

[0002] The Metaverse is a virtual digital world that allows users to interact and experience immersively through different devices and platforms. With the development and popularization of the Metaverse, users have an increasing demand for intelligent interaction with virtual environments.

[0003] According to Chinese patent application number CN202310995469.2, an intelligent interaction method and system of the metaverse are disclosed. First, a user gesture interaction video captured by a camera is obtained. Then, the user gesture interaction video is sampled and feature extracted to obtain a gesture operation type semantic understanding feature vector. Then, based on the gesture operation type semantic understanding feature vector, the user control action intention corresponding to the user gesture interaction video is determined.

[0004] When using some existing intelligent interactive systems, different users may have different habits. Furthermore, the system cannot perform unified analysis of different habits, which will cause the system to be unable to recognize the corresponding interactive information during interaction, and thus cause a decline in user interaction experience. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a metaverse intelligent interaction method and system, which solves the problem that different users may have different interaction methods when interacting, which further solves the problem that the interaction process cannot be well performed and the user interaction experience is reduced.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a metaverse intelligent interaction method, and the method specifically includes the following steps:

[0007] Step 1: The system will collect user interaction information in real time. This data includes but is not limited to voice commands, body movements, facial expressions, etc. Through advanced recognition algorithms, the system processes and analyzes this interaction information and quickly obtains the results of interaction recognition. Specifically, the system will determine the interaction method used by the user based on the content of the user's interaction. This can be voice interaction that conveys instructions through voice, or action interaction that performs operations through body movements.

[0008] Step 2: When the interaction recognition result obtained is a voice interaction result, the voice information of the current user is collected and processed, and further interaction is performed according to the processed result. The specific processing process is as follows:

[0009] The system will capture and record the user's voice input in real time, and then send the voice information to the advanced voice processing process, in which the system will analyze the voice information through accurate voice recognition technology to determine whether it contains executable interactive instructions. After confirming the interactive intention in the voice information, the system will generate the corresponding recognition signal. Specifically, the recognition signal is mainly divided into two categories: real-time interaction signal and voice analysis signal.

[0010] When the generated recognition signal is a real-time interaction signal, the system will interact with the user in real time based on the voice information, generate interaction results, and transmit the interaction results to the current user.

[0011] When the generated recognition signal is a speech analysis signal, the current comprehensive factors of the speech information are analyzed, and the comprehensive factors here specifically include environmental factors and language factors, and the environmental factors and language factors are processed separately.

[0012] The specific ways to deal with environmental factors are:

[0013] The system first captures the sound signals in the environment through the sound collection device. These signals contain the user's language information and other background noise. Then, the system uses the acoustic analysis algorithm to process these sound signals to distinguish and extract the decibel levels of human voice and noise. The measurement of human voice decibel (d1) and noise decibel (d2) allows the system to understand the sound composition in the current environment;

[0014] Next, the system calculates the ratio of noise decibels to human voice decibels (d2 / d1), which is a key indicator for judging speech clarity. The ratio will be compared with a preset threshold, which is set by the operator based on the actual environment and needs. If the ratio is higher than the preset threshold, it means that the noise level is relatively high, which may have a negative impact on the accuracy of speech recognition. The system therefore generates noise impact information, prompting the user to interact in a quieter environment or take measures to reduce background noise;

[0015] On the contrary, if the ratio is lower than the preset threshold, it indicates that the human voice decibel is dominant and the speech information is clear and reliable. In this case, the system will generate secondary analysis information.

[0016] Then, the generated secondary analysis information is processed, and in this processing process, the language in the voice information is mainly analyzed, the language of the voice information is judged, and it is classified into standard voice information and non-standard voice information. The standard voice information here means that the voice information obtained is all in Mandarin version, and the opposite is non-Mandarin version content;

[0017] The generated non-standard voice information is further analyzed, the acquired voice information is split, and a single font is obtained after splitting, and then the single font is recognized and analyzed. The non-standard fonts in the single font are screened and marked, and standard font information is generated at the same time. The marked font information is then displayed to the operator, and a re-entry signal is further generated.

[0018] Further, the current user re-interacts based on the displayed non-standard font information and noise impact information, and the system identifies based on the information obtained from the re-interaction, and performs analysis in the same way.

[0019] For the generated real-time interaction signal, the system displays the corresponding interaction results according to the current user's voice interaction results, and displays the interaction results to the current user.

[0020] Step 3: When the interaction recognition result obtained is an action interaction result, obtain the real-time interaction video of the current user, extract the content of the real-time interaction video, obtain the video content, then draw key points according to the video content to obtain action features, and analyze the action features. The specific analysis method is as follows:

[0021] The real-time interactive video of the current user is obtained, and the video content is extracted through video analysis. At the same time, the action characteristics of the current user are analyzed according to the video content. The specific analysis method is: three-dimensional modeling is performed on the body movements of the current user, and then the joint points of the current user are obtained and recorded as feature points, and the feature points are represented by three-dimensional coordinates;

[0022] Then, the movement characteristics of the feature points within the time period t are obtained, and the movement characteristics are traced to obtain the user trajectory. The specific value of the time period t here is set by the operator. Similarly, the interactive instructions are standardized to obtain the standard trajectory. Then, the standard trajectory is compared and analyzed with the user trajectory, as follows:

[0023] When the standard trajectory is a repeated trajectory, and the repeated trajectory here means that the user needs to perform the same operation according to the corresponding instruction, such as blinking and turning the head, etc., the action duration of the standard trajectory is obtained and recorded as T1, and the number of repetitions of the standard trajectory is obtained and recorded as C1. Then the movement frequency of a single standard trajectory is calculated and recorded as P1. Similarly, the movement frequency of the user trajectory is calculated P2, and the two are compared. When the frequency difference between the two is within the frequency difference range, it means that the user trajectory corresponding to the current user meets the instruction requirements, and a real-time interaction signal is generated. On the contrary, when the frequency difference between the two is not within the frequency difference range, it means that the user trajectory corresponding to the current user does not meet the instruction requirements, and a trajectory analysis signal is generated;

[0024] When the standard trajectory is a single trajectory, and the single trajectory here represents a one-time user action, such as clapping and nodding, the characteristic amplitude of the standard trajectory is analyzed, and the specific characteristic amplitude represents the rotation angle corresponding to the standard trajectory, and the characteristic amplitude is recorded as F1. At the same time, the characteristic amplitude of the user trajectory corresponding to the current user is obtained and recorded as F2, and the difference between F1 and F2 is calculated, and then the difference is compared with the amplitude difference range. When the difference is within the amplitude difference range, it means that the user trajectory corresponding to the current user meets the instruction requirements, and a real-time interaction signal is generated. On the contrary, when the difference between the two is not within the amplitude difference range, it means that the user trajectory corresponding to the current user does not meet the instruction requirements, and a re-entry signal is generated;

[0025] For the generated real-time interaction signal, the system displays the corresponding interaction results according to the current user's action interaction results, and displays the interaction results to the current user.

[0026] Step 4: For the generated trajectory analysis signal, further secondary recognition analysis is performed on the current user's real-time interaction video, and the specific analysis method is as follows:

[0027] The video background in the real-time interactive video is obtained, and the background color in the video background is extracted. The background colors are compared for similarity, and the background colors that are similar to the real-time color of the current user are screened out as similar colors. Then, the state of the similar colors is judged. If the current similar color is static, it means that the video background color has no effect on the interactive recognition, and a re-entry signal is generated. On the contrary, if the current similar color is dynamic, further analysis is performed on whether there is a characteristic action of the similar color.

[0028] When similar colors have corresponding characteristic actions, objects corresponding to similar colors are screened to determine whether the shapes of the objects match. If the shapes of the objects match, the objects are removed, and real-time interaction signals are further generated based on the remaining similar colors. Otherwise, if the shapes of the objects do not match, they are not processed.

[0029] When there is no corresponding characteristic action for a similar color, no processing is performed and a re-entry signal is generated.

[0030] Step 5: The system generates corresponding interaction results based on the real-time interaction signal and displays the interaction results to the current user.

[0031] The present invention provides a metaverse intelligent interaction method and system. Compared with the prior art, it has the following beneficial effects:

[0032] The present invention identifies and analyzes the interactive information generated by the user, classifies the interactive information, and performs separate analysis and identification on voice interaction and action interaction. For voice interaction, the interactive content is identified and analyzed from the two aspects of environment and language to avoid interaction problems due to environment and language reasons. For action interaction, the user is analyzed by acquiring the video of the interaction process, and combined with the analysis of multiple factors, the recognition accuracy of the interaction process is improved, thereby further improving the user interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a block diagram of the system principle of the present invention. DETAILED DESCRIPTION

[0034] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0035] See also Figure 1 The present application provides a method for intelligent interaction of a metaverse, and the method for intelligent interaction comprises the following steps:

[0036] Step 1: The system will collect user interaction information in real time. This data includes but is not limited to voice commands, body movements, facial expressions, etc. Through advanced recognition algorithms, the system processes and analyzes this interaction information and quickly obtains the results of interaction recognition. Specifically, the system will determine the interaction method used by the user based on the content of the user's interaction. This can be voice interaction that conveys instructions through voice, or action interaction that performs operations through body movements.

[0037] For example, when a user says "open the music player", the system captures this instruction through voice recognition technology and determines that the user is using voice interaction. On the other hand, if the user controls the light switch by waving his hand, the system will capture this action through motion recognition technology and determine that the user is using motion interaction.

[0038] The interaction information of the current user is collected, and the interaction information is processed and analyzed using an advanced recognition algorithm to obtain an interaction recognition result, and the interaction recognition result includes a voice interaction result and a motion interaction result.

[0039] Step 2: When the interaction recognition result obtained is a voice interaction result, the voice information of the current user is collected and processed, and further interaction is performed according to the processed result. The specific processing process is as follows:

[0040] The system will capture and record the user's voice input in real time, and then send the voice information to the advanced voice processing process, in which the system will analyze the voice information through accurate voice recognition technology to determine whether it contains executable interactive instructions. After confirming the interactive intention in the voice information, the system will generate the corresponding recognition signal. Specifically, the recognition signal is mainly divided into two categories: real-time interaction signal and voice analysis signal.

[0041] When the generated recognition signal is a real-time interaction signal, the system will interact with the user in real time based on the voice information, generate interaction results, and transmit the interaction results to the current user.

[0042] When the generated recognition signal is a speech analysis signal, the current comprehensive factors of the speech information are analyzed, and the comprehensive factors here specifically include environmental factors and language factors, and the environmental factors and language factors are processed separately.

[0043] The specific ways to deal with environmental factors are:

[0044] The system first captures the sound signals in the environment through the sound collection device. These signals contain the user's language information and other background noise. Then, the system uses the acoustic analysis algorithm to process these sound signals to distinguish and extract the decibel levels of human voice and noise. The measurement of human voice decibel (d1) and noise decibel (d2) allows the system to understand the sound composition in the current environment;

[0045] Next, the system calculates the ratio of noise decibels to human voice decibels (d2 / d1), which is a key indicator for judging speech clarity. The ratio will be compared with a preset threshold, which is set by the operator based on the actual environment and needs. If the ratio is higher than the preset threshold, it means that the noise level is relatively high, which may have a negative impact on the accuracy of speech recognition. The system therefore generates noise impact information, prompting the user to interact in a quieter environment or take measures to reduce background noise;

[0046] On the contrary, if the ratio is lower than the preset threshold, it indicates that the human voice decibel is dominant and the speech information is clear and reliable. In this case, the system will generate secondary analysis information.

[0047] Then, the generated secondary analysis information is processed, and in this processing process, the language in the voice information is mainly analyzed, the language of the voice information is judged, and it is classified into standard voice information and non-standard voice information. The standard voice information here means that the voice information obtained is all in Mandarin version, and the opposite is non-Mandarin version content;

[0048] The generated non-standard voice information is further analyzed, the acquired voice information is split, and a single font is obtained after splitting, and then the single font is recognized and analyzed. The non-standard fonts in the single font are screened and marked, and standard font information is generated at the same time. The marked font information is then displayed to the operator, and a re-entry signal is further generated.

[0049] Further, the current user re-interacts based on the displayed non-standard font information and noise impact information, and the system identifies based on the information obtained from the re-interaction, and performs analysis in the same way.

[0050] For the generated real-time interaction signal, the system displays the corresponding interaction results according to the current user's voice interaction results, and displays the interaction results to the current user.

[0051] Specifically, after receiving the re-entered information, the system will display it to the user who is interacting. The user will then re-enter the information based on the current information prompt, and then interact through system recognition. After the interaction, the corresponding interaction result will be displayed to the current user.

[0052] Step 3: When the interaction recognition result obtained is an action interaction result, obtain the real-time interaction video of the current user, extract the content of the real-time interaction video, obtain the video content, then draw key points according to the video content to obtain action features, and analyze the action features. The specific analysis method is as follows:

[0053] The real-time interactive video of the current user is obtained, and the video content is extracted through video analysis. At the same time, the action characteristics of the current user are analyzed according to the video content. The specific analysis method is: three-dimensional modeling is performed on the body movements of the current user, and then the joint points of the current user are obtained and recorded as feature points, and the feature points are represented by three-dimensional coordinates;

[0054] Then, the movement characteristics of the feature points within the time period t are obtained, and the movement characteristics are traced to obtain the user trajectory. The specific value of the time period t here is set by the operator. Similarly, the interactive instructions are standardized to obtain the standard trajectory. Then, the standard trajectory is compared and analyzed with the user trajectory, as follows:

[0055] When the standard trajectory is a repeated trajectory, and the repeated trajectory here means that the user needs to perform the same operation according to the corresponding instruction, such as blinking and turning the head, etc., the action duration of the standard trajectory is obtained and recorded as T1, and the number of repetitions of the standard trajectory is obtained and recorded as C1. Then the movement frequency of a single standard trajectory is calculated and recorded as P1. Similarly, the movement frequency of the user trajectory is calculated P2, and the two are compared. When the frequency difference between the two is within the frequency difference range, it means that the user trajectory corresponding to the current user meets the instruction requirements, and a real-time interaction signal is generated. On the contrary, when the frequency difference between the two is not within the frequency difference range, it means that the user trajectory corresponding to the current user does not meet the instruction requirements, and a trajectory analysis signal is generated;

[0056] When the standard trajectory is a single trajectory, and the single trajectory here represents a one-time user action, such as clapping and nodding, the characteristic amplitude of the standard trajectory is analyzed, and the specific characteristic amplitude represents the rotation angle corresponding to the standard trajectory, and the characteristic amplitude is recorded as F1. At the same time, the characteristic amplitude of the user trajectory corresponding to the current user is obtained and recorded as F2, and the difference between F1 and F2 is calculated, and then the difference is compared with the amplitude difference range. When the difference is within the amplitude difference range, it means that the user trajectory corresponding to the current user meets the instruction requirements, and a real-time interaction signal is generated. On the contrary, when the difference between the two is not within the amplitude difference range, it means that the user trajectory corresponding to the current user does not meet the instruction requirements, and a re-entry signal is generated;

[0057] For the generated real-time interaction signal, the system displays the corresponding interaction results according to the current user's action interaction results, and displays the interaction results to the current user.

[0058] Step 4: For the generated trajectory analysis signal, further secondary recognition analysis is performed on the current user's real-time interaction video, and the specific analysis method is as follows:

[0059] The video background in the real-time interactive video is obtained, and the background color in the video background is extracted. The background colors are compared for similarity, and the background colors that are similar to the real-time color of the current user are screened out as similar colors. Then, the state of the similar colors is judged. If the current similar color is static, it means that the video background color has no effect on the interactive recognition, and a re-entry signal is generated. On the contrary, if the current similar color is dynamic, further analysis is performed on whether there is a characteristic action of the similar color.

[0060] When similar colors have corresponding characteristic actions, objects corresponding to similar colors are screened to determine whether the shapes of the objects match. If the shapes of the objects match, the objects are removed, and real-time interaction signals are further generated based on the remaining similar colors. Otherwise, if the shapes of the objects do not match, they are not processed.

[0061] When there is no corresponding characteristic action for similar colors, no processing is performed and a re-entry signal is generated. Specifically, when there is a corresponding characteristic action for similar colors, the system will screen the objects corresponding to the similar colors and use a series of graphic recognition technologies to evaluate whether the shape of the object matches the shape in the predefined shape database. If the object shape matches, it means that the dynamic change may be caused by a non-target object. The system will remove the object from consideration and generate a real-time interactive signal based on the remaining similar color information for more in-depth processing. If the object shape does not match, it means that the dynamic change may be caused by irrelevant factors. The system chooses not to process it and generate a re-entry signal.

[0062] For example, suppose that during the user's interaction with the system, the system detects a dynamic red area in the background that is similar in color to the user's clothing. The system first analyzes whether this red area is accompanied by characteristic movements, such as waving or nodding. If so, the system will further analyze whether the subject of the action is the user, such as by identifying whether there is a face accompanying the action. If the subject of the action is confirmed to be the user, the system will respond to the interaction based on the user's action; if not, the system will ignore this dynamic area to avoid misoperation. If the dynamic red area is not accompanied by any characteristic movements, the system will determine that the change may be caused by non-interactive factors such as wind-blown curtains, and prompt the user to perform the interactive action again.

[0063] Step 5: The system generates corresponding interaction results based on the real-time interaction signal and displays the interaction results to the current user.

[0064] Embodiment 2

[0065] As a second embodiment of the present invention, a metaverse intelligent interaction system is provided, which includes: an interaction information collection unit, an information specific recognition unit, a voice interaction recognition unit, an action interaction recognition unit, a secondary recognition and analysis unit, and an interaction result output unit.

[0066] The interaction information collection unit is used to collect the interaction information of the current user and transmit the interaction information to the information specific identification unit.

[0067] The information specific identification unit is used to identify and analyze the acquired interaction information, and to obtain voice interaction results and action interaction results by processing and analyzing the interaction information through advanced recognition algorithms, and to transmit the voice interaction results to the voice interaction identification unit and the action interaction results to the action interaction identification unit.

[0068] The voice interaction recognition unit is used to analyze the acquired voice interaction results, obtain real-time interaction signals and voice analysis signals by processing, recognizing and analyzing the voice information of the current user, and perform comprehensive factor analysis on the voice analysis signal to obtain re-entry signals and interaction results, and transmit the interaction results to the interaction result output unit.

[0069] The action interaction recognition unit is used to analyze the acquired action interaction results, process the user's real-time interaction video to obtain video content, further analyze the user's action characteristics to generate interaction results, re-enter signals and trajectory analysis signals, and transmit the trajectory analysis signals to the secondary recognition and analysis unit.

[0070] The secondary recognition and analysis unit is used to perform secondary recognition on the acquired trajectory analysis signal, analyze and generate a real-time interaction signal and a re-entry signal according to the background color in the real-time interaction video, and generate an interaction result according to the real-time interaction signal.

[0071] The interaction result output unit is used to display the obtained interaction result to the current user.

[0072] As the third embodiment of the present invention, the focus is on combining the implementation processes of the first embodiment and the second embodiment.

[0073] Some of the data in the above formulas are dimensionless and numerically calculated. Meanwhile, the contents not described in detail in this specification belong to the prior art known to those skilled in the art.

[0074] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A method for intelligent interaction of the metaverse, characterized in that: The following steps are involved: Step 1: Collect the interaction information of the current user and classify the interaction information into voice interaction results and action interaction results through advanced recognition algorithms; Step 2: Obtain the voice interaction results, generate real-time interaction signals and voice analysis signals by recognizing the user's voice information, and generate secondary analysis information by performing comprehensive factor analysis on the voice analysis signal. Then, split the secondary analysis information according to the language, filter out non-standard fonts, and generate a re-entry signal. The specific analysis method is as follows: Get the sound decibel of the voice information, classify it into human voice decibel d1 and noise decibel d2, and calculate the ratio of noise decibel to human voice decibel and compare it with the preset value set by the operator; If the ratio is greater than a preset value, noise impact information is generated; If the ratio is less than the preset value, secondary analysis information is generated; Step 3: Obtain the action interaction results, extract the content of the user's real-time interaction video, extract the action features according to the video content, generate the user trajectory, and compare the standard trajectory with the user trajectory to generate a real-time interaction signal and a trajectory analysis signal. The specific processing method is as follows: Obtain the standard trajectory characteristic amplitude F1 corresponding to a single trajectory and the current user trajectory characteristic amplitude F2, calculate the difference between F1 and F2 and compare it with the amplitude difference range, if the difference is within the amplitude difference range, generate a real-time interaction signal, otherwise, generate a re-entry signal; Obtain the quasi-trajectory action duration T1 and the number of repetitions C1 corresponding to the repeated trajectory, calculate the single standard trajectory movement frequency P1 and the user trajectory movement frequency P2, and generate a real-time interaction signal if the frequency difference between P1 and P2 is within the frequency difference range; otherwise, generate a trajectory analysis signal; Step 4: Obtain trajectory analysis signals, and generate real-time interaction signals and re-entry signals by performing similarity comparison on the video background of the real-time interactive video and combining feature actions for analysis; Step 5: Generate interaction results based on real-time interaction signals and display the interaction results to the user.

2. The intelligent interaction method of the metaverse according to claim 1, characterized in that: The specific method of identifying the user's voice information in step 2 to generate a real-time interaction signal and a voice analysis signal is: The voice information of the current user is obtained, and the voice information is recognized by the system to determine whether interaction can be performed based on the voice information. At the same time, a recognition signal is generated, and the recognition signal includes a real-time interaction signal and a voice analysis signal.

3. The intelligent interaction method of the metaverse according to claim 1, characterized in that: The specific method of generating the re-entry signal in step 2 is: Determine the language of the voice information and classify it into standard voice information and non-standard voice information, wherein the standard voice information here means that the acquired voice information is all in Mandarin version, and the opposite is non-Mandarin version; The generated non-standard voice information is further analyzed, the acquired voice information is split, and a single font is obtained after splitting, and then the single font is recognized and analyzed. The non-standard fonts in the single font are screened and marked, and a re-entry signal is generated at the same time.

4. The intelligent interaction method of the metaverse according to claim 1, characterized in that: The specific method of generating the user trajectory in step 3 is: The real-time interactive video of the current user is obtained, and its content is extracted through video analysis to obtain the video content. At the same time, the action characteristics of the current user are analyzed according to the video content, the movement characteristics of the feature points within the time period t are obtained, and the movement characteristics are traced to obtain the user trajectory. Similarly, the interactive instructions are standardized to obtain the standard trajectory.

5. The intelligent interaction method of the metaverse according to claim 1, characterized in that: The specific method of analyzing and generating the real-time interaction signal and the re-entry signal in step 4 in combination with the characteristic action is as follows: The video background in the real-time interactive video is obtained, and the background color in the video background is extracted. The background colors are compared for similarity, and the background colors that are similar to the real-time color of the current user are screened out and recorded as similar colors. Then the state of the similar colors is judged. If the current similar color is static, it means that the video background color has no effect on the interactive recognition, and a re-entry signal is generated. On the contrary, if the current similar color is dynamic, further analysis is performed on whether there is a characteristic action of the similar color.

6. The intelligent interaction method of the metaverse according to claim 5, characterized in that: The specific method of analyzing whether similar colors have characteristic actions in step 4 is: When similar colors have corresponding characteristic actions, objects corresponding to similar colors are screened to determine whether the shapes of the objects match. If the shapes of the objects match, the objects are removed, and real-time interaction signals are further generated based on the remaining similar colors. Otherwise, if the shapes of the objects do not match, they are not processed. When there is no corresponding characteristic action for a similar color, no processing is performed and a re-entry signal is generated.

7. A metaverse intelligent interaction system, used to execute a metaverse intelligent interaction method according to any one of claims 1 to 6, characterized in that: include: Interaction information collection unit, information specific recognition unit, voice interaction recognition unit, action interaction recognition unit, secondary recognition and analysis unit and interaction result output unit; The interaction information collection unit is used to collect the interaction information of the current user and transmit the interaction information to the information specific identification unit; The information specific recognition unit is used to recognize and analyze the acquired interaction information, process and analyze the interaction information through advanced recognition algorithms to obtain voice interaction results and action interaction results, and transmit the voice interaction results to the voice interaction recognition unit, and transmit the action interaction results to the action interaction recognition unit; A voice interaction recognition unit is used to analyze the acquired voice interaction results, obtain a real-time interaction signal and a voice analysis signal by processing, recognizing and analyzing the voice information of the current user, and perform a comprehensive factor analysis on the voice analysis signal to obtain a re-entry signal and an interaction result, and transmit the interaction result to the interaction result output unit; The action interaction recognition unit is used to analyze the obtained action interaction results, process the user's real-time interaction video to obtain video content, further analyze the user's action characteristics to generate interaction results, re-enter signals and trajectory analysis signals, and transmit the trajectory analysis signals to the secondary recognition and analysis unit; A secondary recognition and analysis unit, used for performing secondary recognition on the acquired trajectory analysis signal, analyzing and generating a real-time interaction signal and a re-entry signal according to the background color in the real-time interaction video, and generating an interaction result according to the real-time interaction signal; The interaction result output unit is used to display the obtained interaction result to the current user.

Citation Information

Patent Citations

  • Intelligent Interaction Methods and Systems in the Metaverse

    CN116719419B

  • Metacosm cross-screen virtual-real interaction method and system

    CN117891351A

  • Intelligent voice interaction method and apparatus, device, and storage medium

    WO2024022027A1