Construction Method of VR Content Library Based on Visual Modeling and Semantic Analysis
Through visual modeling and word meaning analysis methods, combined with image data and script information, a three-dimensional model is generated and attribute labels are matched, which solves the problem of inefficiency in the existing technology, and realizes efficient production of virtual reality content and fast matching of prop models.
Patent Information
- Application Number
- CN202111636340.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-12-28
AI Technical Summary
The existing three-dimensional modeling technology is inefficient and time-consuming, and the prop adjustment efficiency is inefficient in virtual content simulation, and the traditional method is time-consuming.
Using visual modeling and word meaning analysis methods, image data is obtained through camera shooting from multiple angles, a three-dimensional model is generated, and the physical attributes and script information of real objects are combined to achieve the matching of the three-dimensional model and attribute labels.
It improves the efficiency of virtual reality content production, reduces labor costs, and realizes efficient matching of prop models and rapid production of virtual reality content.
Smart Images

Figure CN114429519B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of three-dimensional modeling and virtual content simulation, and particularly relates to a method for constructing a VR content library based on visual modeling and semantic analysis. Background Art
[0002] For three-dimensional modeling, there are generally two major types and three methods. The first type is forward modeling, that is, using professional three-dimensional modeling software to manually model; the second type is reverse modeling, that is, reverse calculating a geometric model from a physical object to generate a 3D model, which can be divided into an active method of using a three-dimensional laser scanning device for measurement for modeling and a passive method of modeling based on images or videos.
[0003] The disadvantages of existing three-dimensional modeling technologies are low efficiency and waste of manpower. Using professional three-dimensional modeling software for manual modeling requires constructing from points to edges and from edges to faces, and to obtain a three-dimensional model consistent with the size, texture, etc. of the real scene, it also involves the measurement work of the real scene and specifying its correspondence with the three-dimensional model. This process is very complex, quite time-consuming and laborious, and the efficiency is not high. At the same time, in virtual content simulation, it is necessary to adjust the props that appear in the virtual content according to the script, and the traditional adjustment method is time-consuming and has low efficiency.
[0004] In view of this, there is an urgent need to propose a method for constructing a VR content library based on visual modeling and semantic analysis. Summary of the Invention
[0005] Therefore, the present invention provides a method for constructing a VR content library based on visual modeling and semantic analysis, which matches the required prop models for virtual reality content production in combination with the production requirements of virtual reality content, thereby improving production efficiency and reducing labor costs.
[0006] The method for constructing a VR content library based on visual modeling and semantic analysis of the present invention includes the following steps:
[0007] Step 1: Obtain the image data of a real object, model according to the image data, and then import the obtained three-dimensional model into the VR content library;
[0008] Step 2: Obtain the physical properties of the real object, map the physical properties to the three-dimensional model correspondingly, and at the same time generate an attribute label and an attribute label set based on the three-dimensional model and import them into the VR content library; the physical properties include color, shape, and the scene where the real object is located;
[0009] Step 3: Obtain a script, perform semantic segmentation on the script to extract semantic information, and then retrieve the three-dimensional model based on the semantic information and match it with the three-dimensional model and its attribute labels.
[0010] Further, the specific process of obtaining the image data of the real object in step 1 is as follows:
[0011] The camera takes multiple images of the real object from multiple angles, and the aperture, shutter, ISO, and white balance parameters are the same when shooting at different angles.
[0012] Further, the specific process of modeling according to the image data in step 1 is as follows:
[0013] Step 1.1: Perform Gaussian blur and downsampling on the image data at different scales to obtain a Gaussian pyramid with m groups and n layers of images in each group;
[0014] Step 1.2: Subtract adjacent layers of images in each group of the Gaussian pyramid to obtain a Gaussian difference pyramid;
[0015] Traverse the pixel points in the Gaussian difference pyramid to detect gray extreme points, and use the gray extreme points as candidate key points;
[0016] Calculate the size and direction of the candidate key points to generate a histogram;
[0017] After rotating the coordinate axis with the peak of the histogram as the main direction, select a sub-region in its scale image with the candidate key point as the center, draw the gradient histogram of 8 directions of the sub-region, and obtain the SIFT feature descriptor after sorting along the direction vector;
[0018] Step 1.3: Compare the SIFT feature descriptors of images at different angles, calculate the Euclidean distance between a key point in one image and all key points in another image to obtain the nearest point and the second-nearest point, calculate the ratio of the distance between the key point and the nearest point and the distance between the key point and the second-nearest point. If the ratio is less than the threshold, use the key point and the nearest point as a pair of matching points;
[0019] Step 1.4: After calculating at least 8 pairs of matching points, calculate the fundamental matrix F, and the fundamental matrix includes the internal and external parameters of the camera;
[0020] Step 1.5: Inversely calculate the projection matrix from the image point to the space point through the internal and external parameters of the camera, thereby calculating the spatial coordinates, finally obtaining a three-dimensional point cloud, and triangulating the three-dimensional point cloud to finally obtain a three-dimensional model.
[0021] Further, step 2 specifically includes:
[0022] Step 2.1: Obtain the color, shape, and the scene where the real object is located, and at the same time generate attribute labels S1, S2... S N ;
[0023] Step 2.2: Map the color of the real object to the corresponding 3D model and record it in the attribute label corresponding to the 3D model;
[0024] Step 2.3: Map the shape of the real object to the corresponding 3D model and record it in the attribute label corresponding to the 3D model;
[0025] Step 2.4: Map the scene where the real object is located to the corresponding 3D model and record it in the attribute label corresponding to the 3D model;
[0026] Step 2.5: Import S1, S2... S N into the VR content library.
[0027] Further, step 3 specifically includes:
[0028] Step 3.1: Obtain the script and split the script content sentence by sentence;
[0029] Then, perform word segmentation, part-of-speech tagging, and stop word filtering on the sentences through ICTCLAS;
[0030] Step 3.2: Extract the prop information in the script and establish data labels Y1, Y2... Y for the props N ;
[0031] Step 3.3: Extract the prop color information in the script and record the prop color information in the corresponding data label;
[0032] Step 3.4: Extract the prop shape information in the script and record the prop shape information in the corresponding data label;
[0033] Step 3.5: Extract the prop scene information in the script and record the prop scene information in the corresponding data label;
[0034] Step 3.6: Perform an intersection operation on each data set of prop information with the data set of 3D model attribute labels in the VR content library one by one to achieve the matching of the script and the virtual reality content.
[0035] Further, the extraction of prop information in step 3.2 includes:
[0036] Step 3.2.1: Log common props in the word segmentation dictionary and directly extract them;
[0037] Step 3.2.2: Extract high-frequency nouns, and at the same time retrieve the context of high-frequency nouns including "verb + prop", "verb + quantifier + prop", "adjective + prop", "preposition + prop", and then extract the props.
[0038] Further, the extraction of the prop color information in step 3.3 is specifically as follows:
[0039] Taking the prop vocabulary as the center, retrieve the characters within a predetermined range of its context, search for the vocabulary representing colors, and extract the colors of the props.
[0040] Further, the extraction of the prop shape information in step 3.4 is specifically as follows:
[0041] Taking the prop vocabulary as the center, retrieve the characters within a predetermined range of its context, search for the vocabulary representing shapes, and extract the shapes of the props.
[0042] Furthermore, the extraction of the scene information where the props are located in step 3.5 is specifically as follows:
[0043] Taking the prop vocabulary as the center, retrieve the characters within a predetermined range of its context, search for the vocabulary of "verb + location" and "preposition + location", and extract the scenes where the props are located.
[0044] The above technical solution of the present invention has the following advantages compared with the prior art:
[0045] Combined with the production requirements of virtual reality content, match the required prop models for virtual reality content production, thereby improving production efficiency and reducing labor costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a schematic flowchart provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] Embodiment 1.
[0049] The method for constructing a VR content library based on visual modeling and semantic analysis of the present invention includes the following steps:
[0050] Step 1: Obtain the image data of real objects, model according to the image data, and then import the three-dimensional model into the VR content library;
[0051] Step 2: Obtain the physical properties of real objects, map the physical properties to the three-dimensional model correspondingly, and at the same time generate the attribute labels and attribute label sets based on the three-dimensional model and import them into the VR content library; the physical properties include colors, shapes, and the scenes where the real objects are located;
[0052] Step 3: Obtain a script, perform word sense segmentation on the script to extract word sense information, and then retrieve 3D models based on the word sense information and match them with the 3D models and their attribute tags.
[0053] Furthermore, the specific process of obtaining the image data of the real object in Step 1 is as follows:
[0054] The camera takes multiple images of the real object from multiple angles, and the aperture, shutter speed, ISO, and white balance parameters are the same when shooting at different angles.
[0055] Furthermore, the specific process of modeling based on the image data in Step 1 is as follows:
[0056] Step 1.1: Perform Gaussian blur and downsampling on the image data at different scales to obtain a Gaussian pyramid with m groups and n layers of images in each group;
[0057] Step 1.2: Subtract adjacent layers of images in each group of the Gaussian pyramid to obtain a Gaussian difference pyramid;
[0058] Traverse the pixel points in the Gaussian difference pyramid to detect gray extreme points, and use the gray extreme points as candidate key points;
[0059] Calculate the size and direction of the candidate key points to generate a histogram;
[0060] After rotating the coordinate axis with the peak of the histogram as the main direction, select a sub-region in its scale image with the candidate key point as the center, draw the gradient histogram of 8 directions in the sub-region, and obtain the SIFT feature descriptor after sorting along the direction vector;
[0061] Step 1.3: Compare the SIFT feature descriptors of images at different angles, calculate the Euclidean distance between a key point in one image and all key points in another image to obtain the nearest point and the second nearest point, calculate the ratio of the distance between the key point and the nearest point and the distance between the key point and the second nearest point. If the ratio is less than the threshold, use the key point and the nearest point as a pair of matching points;
[0062] Step 1.4: After calculating at least 8 pairs of matching points, calculate the fundamental matrix F, and the fundamental matrix includes the internal and external parameters of the camera;
[0063] Step 1.5: Calculate the projection matrix from the image point to the space point by inversely calculating the internal and external parameters of the camera, thereby calculating the space coordinates, finally obtaining the 3D point cloud, and triangulating the 3D point cloud to finally obtain the 3D model.
[0064] Furthermore, Step 2 specifically includes:
[0065] Step 2.1: Obtain the color, shape, and the scene where the real object is located, and simultaneously generate the attribute tags S1, S2... S for the corresponding 3D model. N ;
[0066] Step 2.2: Map the color of the real object into the corresponding 3D model and record it in the attribute tag corresponding to the 3D model.
[0067] Step 2.3: Map the shape of the real object into the corresponding 3D model and record it in the attribute tag corresponding to the 3D model.
[0068] Step 2.4: Map the scene where the real object is located into the corresponding 3D model and record it in the attribute tag corresponding to the 3D model.
[0069] Step 2.5: Import S1, S2... S N into the VR content library.
[0070] Furthermore, the specific steps of Step 3 include:
[0071] Step 3.1: Obtain the script and split the script content sentence by sentence.
[0072] Then, perform word segmentation, part-of-speech tagging, and stop word filtering on the sentences through ICTCLAS.
[0073] Step 3.2: Extract the prop information in the script and establish the data tags Y1, Y2... Y for the props. N ;
[0074] Step 3.3: Extract the prop color information in the script and record the prop color information in the corresponding data tag.
[0075] Step 3.4: Extract the prop shape information in the script and record the prop shape information in the corresponding data tag.
[0076] Step 3.5: Extract the prop scene information in the script and record the prop scene information in the corresponding data tag.
[0077] Step 3.6: Perform an intersection operation on each data set of prop information and the data set of 3D model attribute tags in the VR content library one by one to achieve the matching of the script and the virtual reality content.
[0078] Furthermore, the extraction of the prop information in Step 3.2 includes:
[0079] Step 3.2.1: Log the common props in the word segmentation dictionary and directly extract them.
[0080] Step 3.2.2: Extract high-frequency nouns, and at the same time retrieve the context of high-frequency nouns, including "verb + prop", "verb + quantifier + prop", "adjective + prop", "preposition + prop", and then extract the props.
[0081] Furthermore, the specific method for extracting the prop color information in Step 3.3 is as follows:
[0082] Centering on the prop vocabulary, retrieve a predetermined range of characters in its context to search for color-related vocabulary and extract the prop color.
[0083] Furthermore, the specific method for extracting the prop shape information in Step 3.4 is as follows:
[0084] Centering on the prop vocabulary, retrieve a predetermined range of characters in its context to search for shape-related vocabulary and extract the prop shape.
[0085] Furthermore, the specific method for extracting the scene information where the prop is located in Step 3.5 is as follows:
[0086] Centering on the prop vocabulary, retrieve a predetermined range of characters in its context to search for vocabulary of "verb + location" and "preposition + location", and extract the scene where the prop is located.
[0087] Embodiment 2
[0088] This embodiment is a specific embodiment actually used on the basis of Embodiment 1.
[0089] For data collection of real objects, a single-lens reflex camera is used for shooting. Multiple angles (such as top view, front view, upward view, left view, right view, front view, rear view, etc.) should be used for combined photography, and the aperture, shutter speed, ISO, and white balance of the photos should be as consistent as possible.
[0090] Perform Gaussian blur and downsampling (skipping points sampling) on the image at different scales to generate a Gaussian pyramid. The Gaussian pyramid has a total of m groups, and each group has a total of n layers of images. The number of layers n is generally 3 to 5 layers. Convolve the original image with Gaussian functions at different scales, that is, L(x, y, σ) = G(x, y, σ) * I(x, y), σ is the scale space factor that determines the degree of image blur, generally taking 1.6. Among them, G(x, y, σ) is a Gaussian function with a variable convolution kernel, and I(x, y, σ) is the input image.
[0091] For example, first multiply each layer of the images in the 0th group by different scale factors Then perform downsampling on the third-to-last layer image of the 0th group with a scale factor of 2 to obtain the first layer image of the 1st group. Multiply each layer of the images in the 1st group by different scale factors (0, σ, kσ, k2 σ, k 3 σ, k 4 σ, …… k n-2 σ), Repeat the above steps to obtain the Gaussian pyramid.
[0092] Subtract adjacent layers in each group of the Gaussian pyramid to generate the Difference of Gaussian (DoG) pyramid. That is:
[0093] D(x, y, σ) = (G(x, y, kσ) - G(x, y, σ)) * I(x, y) = L(x, y, kσ) - L(x, y, σ)
[0094] Detect extreme points in the Difference of Gaussian pyramid. Compare a pixel with its 8 adjacent pixels at the same scale and 18 adjacent pixels at the corresponding positions in the adjacent scales above and below, for a total of 26 points. When the gray value of this pixel is an extreme value, this point is considered a candidate key point.
[0095] According to
[0096] and
[0097] Calculate the size and direction of the key points, and summarize the information as a histogram. The peak in the histogram is the main direction.
[0098] Rotate the coordinate axes according to the above main direction to the direction of the key point to ensure rotational invariance. Select a sub-region in its scale image with this point as the center. Among them, this region is a window of size 4×4. At the same time, each sub-region includes 4×4 pixel points. Then, summarize the gradients in 8 directions of the sub-region and draw its histogram. Arrange the order of these 8 direction vectors to obtain a 128-dimensional feature vector, that is, the SIFT feature descriptor.
[0099] Perform similarity comparison on the SIFT feature descriptors of two images. Calculate the Euclidean distance between a key point in one image and all key points in the other image to obtain the nearest point and the second nearest point.
[0100] The above calculation process is as follows: W i (W1, W2, W3, …… W 128 ) and R i (R1, R2, R3, …… R 128 ) two groups of SIFT feature vectors, and their Euclidean distance is If the ratio of the nearest distance to the second nearest distance is less than the threshold (set according to the actual situation), the matching is successful.
[0101] For the matching points m and m' in two images from different perspectives, they are represented in homogeneous coordinates as (u, v, 1) and (u', v', 1).
[0102] m' must be located on the epipolar line l' of m in the image I′, and l' = Fm, where F is called the fundamental matrix. If m’ is located on l’, then m’ T Fm = 0,
[0103] can be linearly represented as Uf = 0,
[0104] U = (u’u, u’v, u’, v’u, v’v, v’, u, v, 1),
[0105] f = (F 11 , F 12 , F 13 , F 21 , F 22 , F 23 , F 31 , F 32 , F 33 ) T , and the fundamental matrix F can be obtained with 8 pairs of matching points.
[0106] The fundamental matrix F contains the internal and external parameters of the camera. Given the internal parameter matrices K and K’ of the two cameras, E = K’ T FK. For the corresponding points m and m’ in the two images, m = Rm’ + T, where R is the rotation matrix and T is the translation vector. The cross product of the translation vector is used to obtain the cross product matrix
[0107]
[0108] Multiply both sides of the above equation by mT[T]X, mT[T]XRm’ = mTEm’. E = [T]XR is the essential matrix containing the rotation matrix and the translation matrix. By performing singular value decomposition (SVD) on the essential matrix, the external parameter matrices R and T of the camera can be obtained. The projection matrix from the image points to the spatial points is inversely calculated through the internal and external parameters of the camera, so as to calculate the spatial coordinates, and finally a three-dimensional point cloud is obtained. The three-dimensional point cloud is triangulated to finally obtain a three-dimensional model.
[0109] The second link is to enter the object attribute labels to improve the prop library for virtual reality content production, including the following steps.
[0110] Enter the multi-dimensional attribute information (such as color information, shape information, scene information, etc.) of the three-dimensional model obtained in the first link into the prop library for virtual reality content production.
[0111] Enter the color information color of the three-dimensional model into the prop library for virtual reality content production, such as red, blue, yellow, …….
[0112] Enter the shape information shape of the three-dimensional model into the prop library for virtual reality content production, such as circular, square, triangular, …….
[0113] Enter the scene information "scene" of the 3D model into the virtual reality content production prop library, such as coffee shops, bookstores, bedrooms, etc.
[0114] From the above steps, the attribute label data sets S1, S2, S3,... of each object model can be supplemented into the virtual reality content production prop library.
[0115] The third link is to match the plot with the virtual reality content production prop library, including the following steps.
[0116] Input the script for virtual reality content production, and split the script content sentence by sentence. Then, perform word segmentation, part-of-speech tagging, and stop word filtering on the sentences through ICTCLAS, such as punctuation marks, tones, and personal pronouns. Use word embedding technology to vectorize the processed text, converting the symbolic information of natural language into digital information in vector form for subsequent processing. Perform information extraction on the processed script, mainly obtaining information about props and their attributes.
[0117] Extract information about the props that appear in the script. Most prop names are included in the word segmentation dictionary, and out-of-vocabulary words need to be extracted based on high word frequency and context part-of-speech. Props will appear relatively frequently throughout the script or in some paragraphs of the script. The context part-of-speech of props has forms such as "verb + prop", "verb + quantifier + prop", "adjective + prop", "preposition + prop", etc. Establish corresponding data sets Y1, Y2, Y3,... for the extracted prop one, prop two, prop three, etc.
[0118] Analyze the color information of the props. The color information of the props is mainly extracted based on rules. Search within k characters of the context centered on the prop word (the size of k is determined by the actual situation), and extract the words representing colors such as "red, yellow, blue,..." that appear in the text. Record the color information "color" of prop one into data set Y1, record the color information "color" of prop two into data set Y2, and so on.
[0119] Analyze the shape information of the props. The extraction of the shape information of the props is similar to the extraction of the color information. Search within k characters of the context centered on the prop word (the size of k is determined by the actual situation), and extract the words representing shapes such as "round, square, triangle,..." that appear in the text. Record the shape information "shape" of prop one into data set Y1, record the shape information "shape" of prop two into data set Y2, and so on.
[0120] The scene information of the props is mainly analyzed through scene indicators and word-formation rules. Scene indicators refer to specific indicators indicating the appearance of scene words in the script, such as "come to XX", "walk to XX", "at XX", etc. Record the scene information scene of Prop 1 into the data set Y1, record the scene information scene of Prop 2 into the data set Y2... and so on.
[0121] According to the above analysis of the prop attribute information, data sets of each prop information can be obtained and matched in the prop library for virtual content production. The data sets of each prop information are respectively subjected to intersection operations with the object model attribute label data sets S1, S2, S3,... in the virtual reality content production prop library one by one. The greater the cardinality of these intersections, the higher the matching degree.
[0122] Obviously, the above embodiments are only examples given for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or variations derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for constructing a VR content library based on visual modeling and semantic analysis, characterized in that It includes the following steps: Step 1: Obtain the image data of a real object, and after modeling based on the image data, import the obtained 3D model into the VR content library; Step 2: Obtain the physical properties of the real object, map the physical properties to the 3D model correspondingly, and at the same time generate an attribute label and an attribute label set based on the 3D model and import them into the VR content library; the physical properties include color, shape, and the scene where the real object is located; Step 3: Obtain a script, perform word sense segmentation on the script to extract word sense information, and then retrieve the 3D model based on the word sense information and match it with the 3D model and its attribute labels; The specific process of the said Step 3 includes: Step 3.1: Obtain a script and segment the script content sentence by sentence; Then perform word segmentation, part-of-speech tagging, and stop word filtering on the sentences through ICTCLAS; Step 3.2: Extract the prop information in the script and establish the data tags Y1, Y2... Y of the props N ; Step 3.3: Extract the prop color information in the script and record the prop color information in the corresponding data label; Step 3.4: Extract the prop shape information in the script and record the prop shape information in the corresponding data label; Step 3.5: Extract the prop scene information in the script and record the prop scene information in the corresponding data label; Step 3.6: Perform intersection operations on each data set of prop information with the 3D model attribute label data set in the VR content library one by one to achieve the matching of the script and the virtual reality content.
2. The construction method of a VR content library based on visual modeling and semantic analysis according to claim 1, characterized in that, The specific process of obtaining the image data of the real object in the said Step 1 is: The camera takes multiple images of the real object from multiple angles, and the aperture, shutter, ISO, and white balance parameters are the same when shooting at different angles.
3. According to the method for constructing a VR content library based on visual modeling and word sense analysis described in claim 2, the specific process of modeling based on the image data in Step 1 is: Step 1.1: Perform Gaussian blur and downsampling on the image data at different scales to obtain a Gaussian pyramid with m groups and n layers of images in each group; Step 1.2: Subtract adjacent layers of images in each group of images in the Gaussian pyramid to obtain a Gaussian difference pyramid; Traverse the pixel points in the Gaussian difference pyramid to detect gray extreme points, and use the gray extreme points as candidate key points; Calculate the size and direction of the candidate key points and generate a histogram; After rotating the coordinate axis with the peak of the histogram as the main direction, select a sub-region in the scale image centered on the candidate key point, draw the gradient histogram in 8 directions of the sub-region, and obtain the SIFT feature descriptor after sorting along the direction vector; Step 1.3: Compare the SIFT feature descriptors of images at different angles, calculate the Euclidean distance between a key point in one image and all key points in another image to obtain the nearest point and the second nearest point, calculate the ratio of the distance between the key point and the nearest point and the distance between the key point and the second nearest point, and if the ratio is less than the threshold, use the key point and the nearest point as a pair of matching points; Step 1.4: After calculating at least 8 pairs of matching points, calculate the fundamental matrix F, and the fundamental matrix includes the internal and external parameters of the camera; Step 1.5: Inversely calculate the projection matrix from the image points to the spatial points through the internal and external parameters of the camera, thereby calculating the spatial coordinates, and finally obtaining a three-dimensional point cloud. Triangulate the three-dimensional point cloud to finally obtain a three-dimensional model.
4. The method for constructing a VR content library based on visual modeling and semantic analysis according to claim 1, wherein The specific steps of Step 2 include: Step 2.1: Obtain the color, shape, and the scene where the real object is located, and simultaneously generate the attribute labels S1, S2... S corresponding to the 3D model N ; Step 2.2: Map the color of the real object to the corresponding three-dimensional model and record it in the attribute label corresponding to the three-dimensional model. Step 2.3: Map the shape of the real object to the corresponding three-dimensional model and record it in the attribute label corresponding to the three-dimensional model. Step 2.4: Map the scene where the real object is located to the corresponding three-dimensional model and record it in the attribute label corresponding to the three-dimensional model. Step 2.5, import S1, S2... S N into the VR content library.
5. The method for constructing a VR content library based on visual modeling and semantic analysis according to claim 1, wherein The prop information extracted in Step 3.2 includes: Step 3.2.1: Enter common props into the word segmentation dictionary and directly extract them. Step 3.2.2: Extract high-frequency nouns, and at the same time retrieve the context of high-frequency nouns, including "verb + prop", "verb + quantifier + prop", "adjective + prop", "preposition + prop", and then extract the props.
6. The construction method of the VR content library based on visual modeling and semantic analysis according to claim 1, characterized in that, The specific method for extracting the prop color information in Step 3.3 is: Centering on the prop vocabulary, retrieve the characters within a predetermined range of its context to retrieve the vocabulary representing colors and extract the colors of the props.
7. The method for constructing a VR content library based on visual modeling and semantic analysis according to claim 1, characterized in that The specific method for extracting the prop shape information in Step 3.4 is: Centering on the prop vocabulary, retrieve the characters within a predetermined range of its context to retrieve the vocabulary representing shapes and extract the shapes of the props.
8. The construction method of the VR content library based on visual modeling and semantic analysis according to claim 1, characterized in that, The specific method for extracting the scene information where the props are located in Step 3.5 is: Centering on the prop vocabulary, retrieve the characters within a predetermined range of its context to retrieve the vocabulary of "verb + location" and "preposition + location", and extract the scenes where the props are located.
Citation Information
Patent Citations
Chinese medicine tongue manifestation retrieval method based on image content analysis
CN102426583A
Method and device for generating three-dimensional scene and terminal equipment
CN108961396A