A hotspot marking method and device of a three-dimensional virtual sand table and a storage medium
By using hotspot area generation models in the browser to automatically annotate hotspot areas of the 3D virtual sandbox, the inefficiency problem in existing technologies is solved, achieving efficient and accurate hotspot annotation, reducing server performance consumption, and improving the construction speed of the 3D virtual sandbox.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN IDEAMAKE SOFTWARE TECH CO LTD
- Filing Date
- 2023-02-09
- Publication Date
- 2026-07-21
Smart Images

Figure CN116310252B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic sand table technology, and further to the application of artificial intelligence (AI) in the field of electronic sand table technology, particularly to a hotspot annotation method, device and storage medium for a three-dimensional virtual sand table. Background Technology
[0002] A 3D virtual sand table, also known as a 3D digital sand table or a 3D electronic sand table, is a 3D electronic model built using basic geographic information data, model data, attribute data, and graphic data as its data foundation and various 3D simulation techniques. It is widely used in fields such as urban planning, military exercises, engineering design, agricultural planning, and environmental governance.
[0003] Among existing 3D virtual sand table display technologies, the sequential frame animation display method is a superior 3D virtual sand table display technology because it can display high-precision images and does not have high performance requirements on the running system.
[0004] However, because 3D virtual sand tables are displayed using a sequence of frame animations, and the 3D virtual sand table in the sequence of frame display method uses frame-by-frame animation to show the changes in 3D scene, a typical 3D virtual sand table consists of 31-121 images. The content of the area displayed in each image is not exactly the same. In the process of constructing a 3D virtual sand table, researchers need to link the content of the area displayed in each image with the same content in other images. In this way, when using the 3D virtual sand table later, after marking a certain area in a certain image of the frame animation image sequence, the corresponding area can be marked in other images.
[0005] Furthermore, if the same mark or punctuation is to be added to other frames of a frame animation image sequence, manual marking is usually used. This method involves manually switching the sequence frames and using the human eye to determine the position of the mark for each frame. This consumes a lot of researchers' time and is very tedious and inefficient. Summary of the Invention
[0006] This application provides a method for hotspot annotation in a three-dimensional virtual sandbox, which can improve the efficiency of hotspot area annotation and avoid the waste of time and manpower costs associated with manual annotation.
[0007] In a first aspect, embodiments of this application provide a method for hotspot annotation in a three-dimensional virtual sandbox, the method being applied to a browser webpage, the method comprising:
[0008] Extract one or more hotspot regions from the first frame image, wherein the first frame image is a frame image in the frame animation image sequence of the three-dimensional virtual sandbox in the web page;
[0009] The hotspot region generation model is used to retrieve target regions in each frame of the frame animation image sequence whose similarity to the hotspot region is higher than a preset threshold.
[0010] The hotspot region and the target region with a similarity higher than a preset threshold are marked as the same object;
[0011] The three-dimensional virtual sand table is displayed by playing the frame animation image sequence. During the display of the three-dimensional virtual sand table, the hot spots and target areas corresponding to the same object are highlighted using the same markings.
[0012] The main feature of this application is that it uses artificial intelligence methods in a browser to train a model in real time to identify and locate specified areas of images in a frame animation image sequence. Only one frame of the 3D virtual sandbox needs to be marked, and the remaining frames in the sequence can be marked automatically. This replaces the time-consuming and tedious work of manually marking each frame one by one, which can improve the marking speed and accuracy. Moreover, the above operations are completed in the browser without consuming additional server performance, thus reducing costs and increasing efficiency.
[0013] Specifically, firstly, hotspot regions in the frame animation image sequence of the 3D virtual sand table labeled by the researchers are obtained. Then, the labeled hotspot regions are processed appropriately to save them as separate images, for example, by cropping the image containing the hotspot region.
[0014] Secondly, the image of the hotspot area is input into the hotspot area generation model for retrieval to determine the target area corresponding to the hotspot area in other images of the frame animation image sequence. However, the images in the frame animation image sequence are highly homogeneous, resulting in many target areas similar to the hotspot area. Therefore, the model's ability is used to distinguish the target area, and a preset threshold is set to measure whether the target area is accurate enough.
[0015] Secondly, after obtaining the target region in each frame of the image, the hotspot region and the target region with a similarity higher than a preset threshold are marked as the same object.
[0016] Finally, the 3D virtual sandbox is marked and displayed by playing the sequence of animated frames. During the display of the 3D virtual sandbox, hotspot areas and target areas corresponding to the same object are highlighted using the same markings. This allows the target areas to be intuitively presented to researchers after the results are obtained, facilitating manual inspection of the target areas. It should be noted that setting the 3D virtual sandbox and hotspot area generation model on the server side will not have such an intuitive effect. Furthermore, since the sequence of frame images of the 3D virtual sandbox is quite similar, performing the above process on the server side would not be convenient for researchers to debug the hotspot area generation model during the model training phase.
[0017] In another possible implementation of the first aspect, the step of retrieving target regions in each frame of the frame animation image sequence whose similarity to the hotspot region is higher than a preset threshold through a hotspot region generation model includes:
[0018] Extract the hotspot region to obtain an image corresponding to the hotspot region;
[0019] The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are input into the hotspot area generation model to obtain the initial target area corresponding to the hotspot area in each frame of the frame animation image sequence. The hotspot area generation model is trained based on the image corresponding to the historical hotspot area, the frame animation image sequence of the historical 3D virtual sandbox, and the target area corresponding to the historical hotspot area in each frame of the frame animation image sequence of the historical 3D virtual sandbox. The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are feature data, and the target area corresponding to the hotspot area in any frame of the frame animation image sequence is label data.
[0020] Initial target regions that have a similarity to the hotspot regions higher than a preset threshold are identified as target regions.
[0021] In the process of identifying target regions in each frame of the frame animation image sequence whose similarity to the hotspot region exceeds a preset threshold, a hotspot region generation model is required. This model is trained based on historical image information corresponding to hotspot regions, the frame animation image sequence of the 3D virtual sandbox, and the target regions corresponding to the hotspot regions in any frame of the frame animation image sequence. The image information corresponding to the hotspot regions and the frame animation image sequence of the 3D virtual sandbox are feature data, and the target regions corresponding to the hotspot regions in any frame of the frame animation image sequence are label data.
[0022] The trained hotspot region generation model has image recognition and analysis processing capabilities. Since the application of the three-dimensional virtual sandbox is relatively frequent, hundreds or thousands of hotspot annotation tasks are generally required in a three-dimensional virtual sandbox. Therefore, in the initial stage of this method, the training of the hotspot region generation model is carried out using steps that are basically the same as those provided in this application. The training is considered complete when the annotation success rate of the final hotspot region generation model meets the requirements.
[0023] In this implementation, a model is used instead of manual labor to identify and retrieve marked hotspot areas, which greatly improves the efficiency and accuracy of hotspot area marking, thereby increasing the construction speed of the 3D virtual sand table.
[0024] In another possible implementation of the first aspect, the three-dimensional virtual sandbox is used for display in a browser webpage, and the hotspot area generation model is deployed in the browser webpage.
[0025] Model training, hotspot identification, and localization are all completed in the browser, which does not consume server performance, supports high concurrency, reduces costs and increases efficiency, and has a high degree of freedom, allowing arbitrary point marking; if there is a problem with the target area output by the model, it can be seen very intuitively; if the identification and localization of hotspot areas are done on the server, it will put a lot of pressure on the server when multiple people are operating.
[0026] In another possible implementation of the first aspect, before inputting the image corresponding to the hotspot region and the frame animation image sequence of the 3D virtual sandbox into the hotspot region generation model to obtain the initial target region corresponding to the hotspot region in each frame of the frame animation image sequence, the method further includes:
[0027] Deep recognition is performed on images corresponding to historical hotspot areas and frame animation image sequences of historical 3D virtual sand table to determine the data information of the images corresponding to the historical hotspot areas and the frame animation image sequences of historical 3D virtual sand table, wherein the data information includes grayscale and resolution;
[0028] Based on the data information of the images corresponding to the historical hotspot areas and the frame animation image sequence of the historical 3D virtual sand table, adjust the data information of the images corresponding to the hotspot areas and the frame animation image sequence of the 3D virtual sand table.
[0029] In this embodiment, before inputting the images corresponding to the hotspot areas and the frame animation image sequence of the 3D virtual sandbox into the hotspot area generation model, the image-related data of the images corresponding to the hotspot areas and the frame animation image sequence of the 3D virtual sandbox are adjusted. By aligning the data information of the images and images input into the model with the data information of the training data of the hotspot area generation model, the accuracy of the hotspot area generation model is improved, and the data information is avoided when the frame animation image sequence of the 3D virtual sandbox is saved or run in a browser webpage.
[0030] In another possible implementation of the first aspect, determining the initial target region with a similarity higher than a preset threshold to the hotspot region as the target region includes:
[0031] The frame of the animated image sequence is determined to include at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold.
[0032] The initial target regions with a similarity higher than a preset threshold to the hotspot region are sorted in descending order of their matching degree with the detailed information of the hotspot region, wherein the detailed information includes at least material information, shape information and color information;
[0033] Among the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold, the initial target region with the highest degree of matching with the detailed information of the hotspot region is selected as the target region.
[0034] Since the initial target regions generated in the hotspot region generation model have already undergone feature point similarity filtering within the model, selecting initial target regions that exceed the preset threshold, there may still be multiple initial target regions.
[0035] Specifically, since the hotspot region generation model retrieves target regions from the frame animation image sequence of the 3D virtual sandbox based on the correlation between images, and the correlation between the frame animation image sequence is strong, multiple predicted initial target regions may appear when outputting the prediction results.
[0036] To address the above situation, the at least two initial target regions with a similarity higher than a preset threshold to the hotspot region are compared with the hotspot region in the model input data in terms of detail information matching degree. In this way, the initial target region that is closest to the hotspot region is determined as the target region, thereby further improving the prediction accuracy of the hotspot region generation model.
[0037] In another possible implementation of the first aspect, sorting the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold according to their matching degree of detailed information with the hotspot region from high to low includes:
[0038] Obtain detailed information about the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than a preset threshold;
[0039] The detailed information matching degree between the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold is determined by comparing them with the detailed information of the hotspot region.
[0040] The initial target regions that are more than two times similar to the hotspot region are sorted in descending order of the degree of matching between the detailed information of the hotspot region and the detailed information of the hotspot region.
[0041] In the process of comparing and sorting the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than a preset threshold, the comparison is performed based on the matching degree of detailed information between the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than the preset threshold. The comparison data is the aforementioned detailed information. Since the target region is selected from the images in the frame animation image sequence that constitutes the 3D virtual sandbox, and the hotspot region and the target region refer to the same region in the 3D virtual sandbox, the material information, shape information, and color information in this region are at least similar. Therefore, by comparing the detailed information, the at least two initial target regions whose similarity to the hotspot region is higher than the preset threshold are sorted to select the initial target region that is closest to the hotspot region as the target region, thereby further improving the accuracy of the hotspot region generation model.
[0042] In another possible implementation of the first aspect, after displaying the 3D virtual sandbox by playing the frame animation image sequence, and highlighting the hotspot areas and target areas corresponding to the same object using the same markers during the display of the 3D virtual sandbox, the method further includes:
[0043] The need to obtain training data for the hotspot region generation model;
[0044] The images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model are formatted to meet the training data requirements of the hotspot region generation model.
[0045] Extract the image corresponding to the target region from the adjusted image;
[0046] The extracted image is input into the hotspot region generation model for retraining;
[0047] Replace the original hotspot generation model with the trained hotspot generation model.
[0048] This implementation mainly involves updating the hotspot region generation model. After the model prediction is completed, since it needs to be displayed in a 3D virtual sandbox, the category, attributes, or format of the target region will be partially adjusted. Therefore, the format of the images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model is adjusted, and these are used as training data to train the hotspot region generation model. Finally, the trained hotspot region generation model is transmitted to the browser webpage to replace the original hotspot region generation model.
[0049] Secondly, embodiments of this application provide a hotspot annotation device for a three-dimensional virtual sandbox. The device includes at least a first extraction unit, a retrieval unit, a marking unit, and a playback unit. This hotspot annotation device for the three-dimensional virtual sandbox is used to implement the method described in any embodiment of the first aspect. The first extraction unit, retrieval unit, marking unit, and playback unit are described below:
[0050] The first extraction unit is used to extract one or more hot spots in the first frame image, wherein the first frame image is a frame image in the frame animation image sequence of the three-dimensional virtual sandbox in the web page;
[0051] The retrieval unit is used to retrieve target regions in each frame of the frame animation image sequence whose similarity to the hotspot region is higher than a preset threshold through a hotspot region generation model;
[0052] A marking unit is used to mark the hotspot region and a target region with a similarity to the hotspot region higher than a preset threshold as the same object;
[0053] The playback unit is used to display the three-dimensional virtual sand table by playing the sequence of frame animation images. During the display of the three-dimensional virtual sand table, the hot spots and target areas corresponding to the same object are highlighted using the same markings.
[0054] The main feature of this application is that it uses artificial intelligence methods in a browser to train a model in real time to identify and locate specified areas of images in a frame animation image sequence. Only one frame of the 3D virtual sandbox needs to be marked, and the remaining frames in the sequence can be marked automatically. This replaces the time-consuming and tedious work of manually marking each frame one by one, which can improve the marking speed and accuracy. Moreover, the above operations are completed in the browser without consuming additional server performance, thus reducing costs and increasing efficiency.
[0055] Specifically, firstly, hotspot regions in the frame animation image sequence of the 3D virtual sand table labeled by the researchers are obtained. Then, the labeled hotspot regions are processed appropriately to save them as separate images, for example, by cropping the image containing the hotspot region.
[0056] Secondly, the image of the hotspot area is input into the hotspot area generation model for retrieval to determine the target area corresponding to the hotspot area in other images of the frame animation image sequence. However, the images in the frame animation image sequence are highly homogeneous, resulting in many target areas similar to the hotspot area. Therefore, the model's ability is used to distinguish the target area, and a preset threshold is set to measure whether the target area is accurate enough.
[0057] Secondly, after obtaining the target region in each frame of the image, the hotspot region and the target region with a similarity higher than a preset threshold are marked as the same object.
[0058] Finally, the 3D virtual sandbox is marked and displayed by playing the sequence of animated frames. During the display of the 3D virtual sandbox, hotspot areas and target areas corresponding to the same object are highlighted using the same markings. This allows the target areas to be intuitively presented to researchers after the results are obtained, facilitating manual inspection of the target areas. It should be noted that setting the 3D virtual sandbox and hotspot area generation model on the server side will not have such an intuitive effect. Furthermore, since the sequence of frame images of the 3D virtual sandbox is quite similar, performing the above process on the server side would not be convenient for researchers to debug the hotspot area generation model during the model training phase.
[0059] In another possible implementation of the second aspect, the retrieval unit is specifically used for:
[0060] Extract the hotspot region to obtain an image corresponding to the hotspot region;
[0061] The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are input into the hotspot area generation model to obtain the initial target area corresponding to the hotspot area in each frame of the frame animation image sequence. The hotspot area generation model is trained based on the image corresponding to the historical hotspot area, the frame animation image sequence of the historical 3D virtual sandbox, and the target area corresponding to the historical hotspot area in each frame of the frame animation image sequence of the historical 3D virtual sandbox. The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are feature data, and the target area corresponding to the hotspot area in any frame of the frame animation image sequence is label data.
[0062] Initial target regions that have a similarity to the hotspot regions higher than a preset threshold are identified as target regions.
[0063] In the process of identifying target regions in each frame of the frame animation image sequence whose similarity to the hotspot region exceeds a preset threshold, a hotspot region generation model is required. This model is trained based on historical image information corresponding to hotspot regions, the frame animation image sequence of the 3D virtual sandbox, and the target regions corresponding to the hotspot regions in any frame of the frame animation image sequence. The image information corresponding to the hotspot regions and the frame animation image sequence of the 3D virtual sandbox are feature data, and the target regions corresponding to the hotspot regions in any frame of the frame animation image sequence are label data.
[0064] The trained hotspot region generation model has image recognition and analysis processing capabilities. Since the application of the three-dimensional virtual sandbox is relatively frequent, hundreds or thousands of hotspot annotation tasks are generally required in a three-dimensional virtual sandbox. Therefore, in the initial stage of this method, the training of the hotspot region generation model is carried out using steps that are basically the same as those provided in this application. The training is considered complete when the annotation success rate of the final hotspot region generation model meets the requirements.
[0065] In this implementation, a model is used instead of manual labor to identify and retrieve marked hotspot areas, which greatly improves the efficiency and accuracy of hotspot area marking, thereby increasing the construction speed of the 3D virtual sand table.
[0066] In yet another possible implementation of the second aspect, the device further includes:
[0067] The identification unit is used to perform depth recognition on images corresponding to historical hotspot areas and frame animation image sequences of historical 3D virtual sand table, and determine the data information of the images corresponding to the historical hotspot areas and the frame animation image sequences of historical 3D virtual sand table, wherein the data information includes grayscale and resolution.
[0068] The adjustment unit is used to adjust the data information of the images corresponding to the historical hotspot areas and the frame animation image sequence of the historical 3D virtual sand table based on the data information of the images corresponding to the historical hotspot areas and the frame animation image sequence of the historical 3D virtual sand table.
[0069] In this embodiment, before inputting the images corresponding to the hotspot areas and the frame animation image sequence of the 3D virtual sandbox into the hotspot area generation model, the image-related data of the images corresponding to the hotspot areas and the frame animation image sequence of the 3D virtual sandbox are adjusted. By aligning the data information of the images and images input into the model with the data information of the training data of the hotspot area generation model, the accuracy of the hotspot area generation model is improved, and the data information is avoided when the frame animation image sequence of the 3D virtual sandbox is saved or run in a browser webpage.
[0070] In another possible implementation of the second aspect, in determining the initial target region whose similarity to the hotspot region is higher than a preset threshold as the target region, the retrieval unit is specifically used for:
[0071] The frame of the animated image sequence is determined to include at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold.
[0072] The initial target regions with a similarity higher than a preset threshold to the hotspot region are sorted in descending order of their matching degree with the detailed information of the hotspot region, wherein the detailed information includes at least material information, shape information and color information;
[0073] Among the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold, the initial target region with the highest degree of matching with the detailed information of the hotspot region is selected as the target region.
[0074] Since the initial target regions generated in the hotspot region generation model have already undergone feature point similarity filtering within the model, selecting initial target regions that exceed the preset threshold, there may still be multiple initial target regions.
[0075] Specifically, since the hotspot region generation model retrieves target regions from the frame animation image sequence of the 3D virtual sandbox based on the correlation between images, and the correlation between the frame animation image sequence is strong, multiple predicted initial target regions may appear when outputting the prediction results.
[0076] To address the above situation, the at least two initial target regions with a similarity higher than a preset threshold to the hotspot region are compared with the hotspot region in the model input data in terms of detail information matching degree. In this way, the initial target region that is closest to the hotspot region is determined as the target region, thereby further improving the prediction accuracy of the hotspot region generation model.
[0077] In another possible implementation of the second aspect, in the aspect of sorting the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold according to their matching degree of detailed information with the hotspot region from high to low, the retrieval unit is specifically used for:
[0078] Obtain detailed information about the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than a preset threshold;
[0079] The detailed information matching degree between the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold is determined by comparing them with the detailed information of the hotspot region.
[0080] The initial target regions that are more than two times similar to the hotspot region are sorted in descending order of the degree of matching between the detailed information of the hotspot region and the detailed information of the hotspot region.
[0081] In the process of comparing and sorting the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than a preset threshold, the comparison is performed based on the matching degree of detailed information between the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than the preset threshold. The comparison data is the aforementioned detailed information. Since the target region is selected from the images in the frame animation image sequence that constitutes the 3D virtual sandbox, and the hotspot region and the target region refer to the same region in the 3D virtual sandbox, the material information, shape information, and color information in this region are at least similar. Therefore, by comparing the detailed information, the at least two initial target regions whose similarity to the hotspot region is higher than the preset threshold are sorted to select the initial target region that is closest to the hotspot region as the target region, thereby further improving the accuracy of the hotspot region generation model.
[0082] In yet another possible implementation of the second aspect, the device further includes:
[0083] The acquisition unit is used to acquire the training data required for the hotspot region generation model;
[0084] The format adjustment unit is used to adjust the format of the images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model, so as to meet the requirements of the training data of the hotspot region generation model.
[0085] The second extraction unit is used to extract the image corresponding to the target region in the adjusted image;
[0086] The training unit is used to input the extracted image into the hotspot region generation model for retraining;
[0087] The replacement unit is used to replace the original hotspot region generation model with the trained hotspot region generation model.
[0088] This implementation mainly involves updating the hotspot region generation model. After the model prediction is completed, since it needs to be displayed in a 3D virtual sandbox, the category, attributes, or format of the target region will be partially adjusted. Therefore, the format of the images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model is adjusted, and these are used as training data to train the hotspot region generation model. Finally, the trained hotspot region generation model is transmitted to the browser webpage to replace the original hotspot region generation model.
[0089] Thirdly, embodiments of this application provide a hotspot labeling device for a three-dimensional virtual sandbox. The hotspot labeling device for the three-dimensional virtual sandbox includes a processor, a memory, and a communication interface. The memory stores a computer program. When the processor executes the computer program, the communication interface is used to send and / or receive data. The hotspot labeling device for the three-dimensional virtual sandbox can execute the method described in the first aspect or any possible implementation of the first aspect.
[0090] It should be noted that the processor included in the hotspot annotation device of the 3D virtual sandbox described in the third aspect above can be a processor specifically designed to execute these methods (referred to as a dedicated processor for distinction), or a processor that executes these methods by calling a computer program, such as a general-purpose processor. Optionally, at least one processor may include both dedicated and general-purpose processors.
[0091] Optionally, the aforementioned computer program can be stored in memory. For example, the memory can be a non-transitory memory, such as read-only memory (ROM), which can be integrated with the processor on the same device or disposed on different devices. This application does not limit the type of memory or the arrangement of the memory and processor.
[0092] In one possible implementation, the at least one memory is located outside the hotspot labeling device of the three-dimensional virtual sandbox.
[0093] In another possible implementation, the at least one memory is located within the hotspot labeling device of the three-dimensional virtual sandbox.
[0094] In another possible implementation, a portion of the memory of the at least one memory is located within the hotspot marking device of the three-dimensional virtual sandbox, while another portion of the memory is located outside the hotspot marking device of the three-dimensional virtual sandbox.
[0095] In this application, the processor and memory may also be integrated into a single device, that is, the processor and memory can be integrated together.
[0096] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed on at least one processor, implements the method described in the first aspect or any of the optional solutions of the first aspect.
[0097] Fifthly, this application provides a computer program product comprising a computer program that, when run on at least one processor, implements the method described in the first aspect or any of the optional solutions of the first aspect.
[0098] Optionally, the computer program product can be a software installation package, which can be downloaded and executed on a computing device when the aforementioned method is required.
[0099] The beneficial effects of the technical solutions provided in the third to fifth aspects of this application can be referred to the beneficial effects of the technical solutions in the first and second aspects, and will not be repeated here. Attached Figure Description
[0100] The accompanying drawings used in the description of the embodiments will be briefly introduced below.
[0101] Figure 1 This is a schematic diagram of the architecture of an electronic device provided in an embodiment of this application;
[0102] Figure 2 This is a flowchart illustrating a hotspot annotation method for a three-dimensional virtual sandbox provided in an embodiment of this application;
[0103] Figure 3 This is a schematic diagram of a hotspot annotation scene in a three-dimensional virtual sandbox provided in an embodiment of this application;
[0104] Figure 4 This is a flowchart illustrating a hotspot region generation model retrieval method provided in an embodiment of this application;
[0105] Figure 5 This is a schematic diagram of the structure of a hotspot marking device for a three-dimensional virtual sandbox provided in an embodiment of this application;
[0106] Figure 6 This is a schematic diagram of the structure of a hotspot annotation device for a three-dimensional virtual sandbox provided in an embodiment of this application. Detailed Implementation
[0107] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0108] The terms "first," "second," "third," and "fourth," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.
[0109] The system architecture used in the embodiments of this application is described below. It should be noted that the system architecture and business scenarios described in this application are for the purpose of more clearly illustrating the technical solutions of this application, and do not constitute a limitation on the technical solutions provided in this application. As those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in this application are also applicable to similar technical problems.
[0110] Please see Figure 1 , Figure 1 This is a schematic diagram of the architecture of an electronic device provided in an embodiment of this application. The architecture mainly includes a browser webpage 10. The electronic device provided in this embodiment of the application mainly executes the method provided in this embodiment of the application by running a computer program in a computer-readable storage medium. The browser webpage includes a three-dimensional virtual sandbox 101 and a hotspot area generation model 102. The three-dimensional virtual sandbox 101 is displayed in the browser webpage 10, and the hotspot area generation model 102 is constructed and trained in the browser webpage 10.
[0111] The 3D virtual sandbox 101 includes a sequence of frame-animated images. This sequence can be obtained by rendering a 3D virtual sandbox scene created using 3D software such as 3D Studio Max (hereinafter referred to as 3d Max), or by obtaining continuous scene images through other means (e.g., creating a series of continuous images using drawing software or taking a series of continuous scene photographs with a camera). The frame-animated images are displayed on a browser webpage 10. The 3D virtual sandbox 101 using the frame-sequence display method employs frame-by-frame animation to showcase the transformation of the 3D scene. For example, a 3D virtual sandbox 101 consists of 31-121 images.
[0112] It should be noted that in this embodiment, the 3D virtual sandbox 101 is constructed and run within a browser webpage 10, and the images of the sequence frames constituting the 3D virtual sandbox 101 are generally stored on a storage server or in the cloud. Optionally, the 3D virtual sandbox 101 can be directly displayed to the user through a terminal display, which can be a common computer screen, tablet screen, projector, 3D imaging device, or other display devices.
[0113] Furthermore, in this embodiment, the three-dimensional virtual sandbox 101 is mainly used to receive user annotations on the first frame image, generate corresponding hotspot areas, and display annotations of hotspot areas and target areas.
[0114] In this embodiment, the hotspot region generation model 102 is implemented based on Google TensorFlow using JavaScript, so that the hotspot region generation model 102 can be constructed and trained in a browser webpage. The hotspot region generation model 102 includes an input layer, an intermediate layer, and an output layer.
[0115] The input layer is mainly used to receive one or more hot spots from the first frame image in the 3D virtual sandbox 101; the intermediate layer is mainly used to retrieve the target area corresponding to the hot spot area in the sequence of frame images constituting the 3D virtual sandbox 101; the output layer is mainly used to transmit the target area retrieved by the intermediate layer to the 3D virtual sandbox 101 so that the 3D virtual sandbox 101 can highlight the hot spot area and the target area.
[0116] Please see Figure 2 , Figure 2 This is a flowchart illustrating a hotspot annotation method for a 3D virtual sandbox provided in an embodiment of this application. This hotspot annotation method is used in a web browser and can be based on... Figure 1 The schematic diagram of the electronic device shown can also be implemented based on other architectures. The method includes, but is not limited to, the following steps:
[0117] Step 201: Extract one or more hotspot regions from the first frame image.
[0118] The frame animation image sequence of a 3D virtual sand table can be obtained by rendering a 3D virtual sand table scene after it has been created using 3D software such as 3D Studio Max (hereinafter referred to as 3dMax), or by obtaining a series of scene images through other means (such as creating a series of continuous images using drawing software or taking a series of continuous scene photos with a camera).
[0119] The first frame image is a frame image in the frame animation image sequence of the three-dimensional virtual sandbox in the web page; the hotspot is the selection point input by the user into the first frame image, and the hotspot area is generated by extending the hotspot input by the user, which can be rectangular or generated based on the boundary of the area selected by the hotspot.
[0120] In this step, hotspots input from the user are received, and hotspot areas are generated. See details in [link to relevant documentation]. Figure 3 The Figure 3 This is a schematic diagram of a scene with hotspot annotations in a three-dimensional virtual sandbox.
[0121] exist Figure 3 In the image, the area selected by the rectangle is the hotspot area, the points in the hotspot area are hotspots, and the area covered by the hotspot area is the area that the user wants to select.
[0122] Step 202: Retrieve target regions in each frame of the frame animation image sequence whose similarity to the hotspot region is higher than a preset threshold using a hotspot region generation model.
[0123] The hotspot generation model is implemented using Google TensorFlow based on JavaScript, which facilitates the creation and training of neural networks within the browser framework. This technology enables model training, hotspot identification, and localization, achieving high concurrency, cost reduction, and efficiency improvement. It also offers a high degree of freedom, allowing arbitrary point marking, and clearly identifies any issues with the target area output by the model. In contrast, placing hotspot identification and localization on a server would place significant pressure on the server under multiple user operations, and improper operation could actually increase the time required.
[0124] In one possible implementation, please refer to [link to relevant documentation] for a detailed implementation of step 202. Figure 4 , Figure 4 The flowchart of a hotspot region generation model retrieval method is shown below:
[0125] Step 401: Extract the hotspot region to obtain the image corresponding to the hotspot region.
[0126] Since the hotspot region generation model is mainly for image recognition and processing, and most regions among the sequence frames of the three-dimensional virtual sandbox are highly homogeneous, before applying the hotspot region generation model, the hotspot regions in the first frame image are extracted to obtain the image corresponding to the hotspot regions. The image corresponding to the hotspot regions is used as input to the hotspot region generation model.
[0127] Optionally, the hotspot area can be extracted by cropping a portion of the image that is framed within the hotspot area.
[0128] Step 402: Input the image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox into the hotspot area generation model to obtain the initial target area corresponding to the hotspot area in each frame of the frame animation image sequence.
[0129] The hotspot region generation model is trained based on images corresponding to historical hotspot regions, a sequence of frame animation images of a historical 3D virtual sandbox, and target regions corresponding to the historical hotspot regions on each frame of the sequence of frame animation images of the historical 3D virtual sandbox. The images corresponding to the hotspot regions and the sequence of frame animation images of the 3D virtual sandbox are feature data, and the target regions corresponding to the hotspot regions on any frame of the sequence of frame animation images are label data.
[0130] Optionally, the code for the hotspot region generation model is constructed by adaptively adjusting the code shown below:
[0131]
[0132] It should be noted that in practical applications, after the model is dynamically trained based on hotspot areas, a JSON (JavaScript Object Notation) data generated based on the front-end architecture will be used to store the model data. This JSON data is mainly used to build a hotspot area generation model in the browser webpage. Based on this trained model data, the selected hotspot area is identified and located by looping through all the remaining sequence frames of images, thus obtaining the initial target area.
[0133] In one optional implementation, before inputting the image corresponding to the hotspot region and the frame animation image sequence of the 3D virtual sandbox into the hotspot region generation model to obtain the initial target region corresponding to the hotspot region in each frame of the frame animation image sequence, the current input data is aligned with the data information corresponding to the historical training data, as follows:
[0134] Deep recognition is performed on images corresponding to historical hotspot areas and frame animation image sequences of historical 3D virtual sandboxes to determine the data information of these images and sequences. This deep recognition is based on deep learning in image recognition, and can be a self-built deep learning neural network model or an existing deep learning neural network model, such as a convolutional neural network, residual network, or residual shrinking network. The data information includes grayscale and resolution. Based on the data information of the images corresponding to the historical hotspot areas and the frame animation image sequences of the historical 3D virtual sandboxes, the data information of these images and sequences is adjusted to ensure that the images and sequences input to the hotspot area generation model are aligned with the standards and have consistent attributes of the training data for the hotspot area generation model, thereby improving the model's prediction accuracy.
[0135] Step 403: Determine the initial target region whose similarity to the hotspot region is higher than a preset threshold as the target region.
[0136] The initial target region and the hotspot region are compared for feature point similarity. Optionally, the feature point similarity comparison is based on the pixel depth values of all pixels in both regions. Pixel depth refers to the number of bits required to store each pixel.
[0137] If the pixel depth values of the initial target region and the hotspot region are similar to a preset threshold, then the similarity between the initial target region and the hotspot region is higher than the preset threshold.
[0138] Given that there may be multiple initial target regions in a single image whose similarity to the hotspot region exceeds a preset threshold, one optional implementation involves filtering these regions to select the initial target region that is closest to the hotspot region. Specifically:
[0139] First, it is determined that a frame of the frame animation image sequence includes at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold;
[0140] Secondly, the at least two initial target regions with a similarity higher than a preset threshold to the hotspot region are sorted in descending order of their matching degree with the detailed information of the hotspot region, wherein the detailed information includes at least material information, shape information and color information;
[0141] Optionally, during the sorting process, detailed information of the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than a preset threshold is selectively obtained. This detailed information includes at least material information, shape information, and color information. The detailed information of the at least two initial target regions whose similarity to the hotspot region is higher than the preset threshold is compared with that of the hotspot region to determine the matching degree of detailed information between the at least two initial target regions whose similarity to the hotspot region is higher than the preset threshold and that of the hotspot region. During this comparison, the detailed information of the at least two initial target regions whose similarity to the hotspot region is higher than the preset threshold is displayed and compared item by item with that of the hotspot region. Table 1 is used as an example for illustration.
[0142] Table 1
[0143]
[0144] Table 1 above mainly displays the item details in the detailed information of the first initial target region, the second initial target region, and the third initial target region and the hotspot region. As can be seen from Table 1, the detailed information of the first initial target region, the second initial target region, and the third initial target region is not completely identical to that of the hotspot region. Therefore, a comparison is performed item by item. In the material information column, the first initial target region and the third initial target region match the hotspot region. Therefore, the similarity score of the material information of the first initial target region and the third initial target region is 2 points. The material information of the second initial target region does not match that of the hotspot region. Therefore, the matching score of the material information of the second initial target region is 0 points. It should be noted that the material information is a dataset. Its main function is to provide data and lighting algorithms for the renderer. The material contains a map, and the map contains a texture. Therefore, the texture type is a part of the material information. Depending on the purpose, the texture will also be divided into different types, such as Diffuse Map, Specular Map, Normal Map, and G loss Map. Another important part of the material information is the lighting model shader, which is used to achieve different rendering effects. Therefore, the material information is quite complex and will not be described in detail in the embodiments of this application. For specific material information, please refer to [link / reference needed]. Figure 3 The schematic diagram of the 3D virtual sandbox shown illustrates the textures, patterns, and lighting.
[0145] In the shape information section, the first initial target area and the second initial target area match the hotspot area. Therefore, the shape information matching degree between the first initial target area, the second initial target area and the hotspot area is scored as 2 points, and the third initial target area is scored as 0 points. It should be noted that since the hotspot area or the initial target area includes a large number of shapes, the shape information is a combination of the shape distribution and quantity of the hotspot area and the initial target area.
[0146] In the color information section, the first and third initial target areas are relatively consistent with the hotspot areas, but there are some differences. Therefore, the color information matching degree between the first and third initial target areas and the hotspot areas is calculated to be 1 point, while the second initial target area is scored 0 points.
[0147] In summary, without setting weights, the first, second, and third initial target regions scored 5, 2, and 3 points respectively. Therefore, the first initial target region had the highest matching degree with the hotspot region's detailed information, followed by the third initial target region, and the second initial target region had the lowest. Thus, the first, second, and third initial target regions were sorted in descending order of their comprehensive scores on the matching degree between the detailed information of the first, second, and third initial target regions and the detailed information of the hotspot region.
[0148] Finally, among the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold, the initial target region with the highest degree of matching with the detailed information of the hotspot region is selected as the target region.
[0149] Step 203: Mark the hotspot region and the target region with a similarity higher than a preset threshold as the same object.
[0150] After each frame in the frame animation image sequence that constitutes the 3D virtual sandbox has a hot spot area or a target area, the hot spot area and the target area are linked together to generate a hot spot area sequence, and the hot spot area matches the frame animation image sequence.
[0151] If the first frame image contains multiple hotspot regions, then the corresponding generated target region will generate different hotspot region sequences.
[0152] Step 204: Display the 3D virtual sand table by playing the frame animation image sequence. During the display of the 3D virtual sand table, the hot spots and target areas corresponding to the same object are highlighted using the same markings.
[0153] The system displays a 3D virtual sandbox by playing a sequence of frame-animated images on Android, iOS, Mac OS, or tv OS platforms. In other words, it displays the images in the frame-animated image sequence one frame at a time to showcase the 3D virtual sandbox.
[0154] Based on the selection pattern of the hotspot region in the first frame image, the selection pattern of the target region in the remaining images is set so that the observer can clearly see that the hotspot region and the target region are the same region. Optionally, the selection pattern includes rectangular selection and boundary selection, wherein the rectangular selection is determined by... Figure 3 As can be seen, in order to clearly mark the hot spot area with a rectangular border, the boundary selection is generated based on the boundary of the content of the selected area of the hot spot area or target area. For example, if the hot spot area is a basketball court in a 3D virtual sandbox, the boundary selection is to highlight the boundary of the basketball court or perform operations such as color change and brightness adjustment to highlight the basketball court.
[0155] In a 3D virtual sandbox, the positions of hotspots used to generate hotspot areas are marked in each frame of the image. In one optional implementation, after obtaining the hotspot area or target area in each frame, the hotspot area or target area is reconstructed. During the reconstruction process, the coordinates of the hotspot in the target area are determined based on the coordinates of the hotspot in the hotspot area. Based on these coordinates, a hotspot view is overlaid at the corresponding position on each frame, completing the dynamic marking of the hotspots. This hotspot view can be, for example, an icon such as an arrow or a button, or text prompts.
[0156] During the display of the 3D virtual sandbox, the hot spots and target areas corresponding to the same object are highlighted using the same markings, so that users can intuitively observe whether there are any problems with the markings of the hot spots or target areas in each frame of the image.
[0157] After highlighting the hotspot areas and target areas corresponding to the same object using the same marker, the hotspot area generation model is self-updated, as follows:
[0158] The requirement for obtaining training data for the hotspot region generation model is a format requirement. According to the format requirement, the training data of the current hotspot region generation model can be aligned with the format of historical training data to improve the performance of the hotspot region generation model.
[0159] The images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model are formatted to meet the training data requirements of the hotspot region generation model.
[0160] Extract the image corresponding to the target region from the adjusted image;
[0161] The extracted image is input into the hotspot region generation model for retraining;
[0162] Replace the original hotspot generation model with the trained hotspot generation model.
[0163] It should be noted that the hotspot region generation model is represented by the GraphDef protocol buffer. Therefore, the updated model is obtained through an RPC interface (e.g., gRPC), which standardizes the protocol buffer processing and uses the RPC interface to create a new TensorFlow session in the Android application.
[0164] In summary, the method provided in this embodiment applies a hotspot region generation model to process hotspot regions. Based on the image corresponding to the hotspot region, it iterates through all remaining sequence frame images to identify and locate the hotspot region. The region with the highest matching rate is selected as the target region from all results, and both the hotspot region and the target region are highlighted. This avoids the hassle of manually selecting hotspot regions and improves the efficiency of hotspot region annotation. Furthermore, the display of the 3D virtual sandbox, the construction and training of the hotspot region generation model are all performed in a browser webpage, allowing users to intuitively observe the annotation status of the hotspot and target regions, avoiding the load cost caused by model construction on a server.
[0165] The methods of the embodiments of this application have been described in detail above, and the apparatus of the embodiments of this application is provided below.
[0166] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of a hotspot marking device 50 for a three-dimensional virtual sandbox provided in this application embodiment. The hotspot marking device 50 for the three-dimensional virtual sandbox can be the aforementioned electronic device or a component within an electronic device. The device 50 may include a first extraction unit 501, a retrieval unit 502, a marking unit 503, and a playback unit 504, wherein each unit is described in detail below.
[0167] The first extraction unit 501 is used to extract one or more hot spots in the first frame image, wherein the first frame image is a frame image in the frame animation image sequence of the three-dimensional virtual sandbox in the web page;
[0168] The retrieval unit 502 is used to retrieve, through a hotspot region generation model, a target region in each frame of the frame animation image sequence whose similarity to the hotspot region is higher than a preset threshold.
[0169] The marking unit 503 is used to mark the hotspot region and the target region with a similarity higher than a preset threshold as the same object;
[0170] The playback unit 504 is used to display the three-dimensional virtual sand table by playing the frame animation image sequence. During the display of the three-dimensional virtual sand table, the hot spot area and target area corresponding to the same object are highlighted using the same marking.
[0171] In one possible implementation, the retrieval unit 502 is specifically used for:
[0172] Extract the hotspot region to obtain an image corresponding to the hotspot region;
[0173] The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are input into the hotspot area generation model to obtain the initial target area corresponding to the hotspot area in each frame of the frame animation image sequence. The hotspot area generation model is trained based on the image corresponding to the historical hotspot area, the frame animation image sequence of the historical 3D virtual sandbox, and the target area corresponding to the historical hotspot area in each frame of the frame animation image sequence of the historical 3D virtual sandbox. The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are feature data, and the target area corresponding to the hotspot area in any frame of the frame animation image sequence is label data.
[0174] Initial target regions that have a similarity to the hotspot regions higher than a preset threshold are identified as target regions.
[0175] In one possible implementation, the device 50 further includes:
[0176] The identification unit is used to perform depth recognition on images corresponding to historical hotspot areas and frame animation image sequences of historical 3D virtual sand table, and determine the data information of the images corresponding to the historical hotspot areas and the frame animation image sequences of historical 3D virtual sand table, wherein the data information includes grayscale and resolution.
[0177] The adjustment unit is used to adjust the data information of the images corresponding to the historical hotspot areas and the frame animation image sequence of the historical 3D virtual sand table based on the data information of the images corresponding to the historical hotspot areas and the frame animation image sequence of the historical 3D virtual sand table.
[0178] In one possible implementation, the retrieval unit 502 is specifically configured to: determine initial target regions whose similarity to the hotspot region is higher than a preset threshold as target regions;
[0179] The frame of the animated image sequence is determined to include at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold.
[0180] The initial target regions with a similarity higher than a preset threshold to the hotspot region are sorted in descending order of their matching degree with the detailed information of the hotspot region, wherein the detailed information includes at least material information, shape information and color information;
[0181] Among the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold, the initial target region with the highest degree of matching with the detailed information of the hotspot region is selected as the target region.
[0182] In one possible implementation, the retrieval unit 502 is specifically used to sort the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold according to their matching degree of detailed information with the hotspot region from high to low.
[0183] Obtain detailed information about the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than a preset threshold;
[0184] The detailed information matching degree between the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold is determined by comparing them with the detailed information of the hotspot region.
[0185] The initial target regions that are more than two times similar to the hotspot region are sorted in descending order of the degree of matching between the detailed information of the hotspot region and the detailed information of the hotspot region.
[0186] In one possible implementation, the device 50 further includes:
[0187] The format adjustment unit is used to adjust the format of the images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model according to the requirements of the training data of the hotspot region generation model, so that the adjusted images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model meet the requirements of the training data of the hotspot region generation model.
[0188] The acquisition unit is used to acquire the training data required for the hotspot region generation model;
[0189] The format adjustment unit is used to adjust the format of the images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model, so as to meet the requirements of the training data of the hotspot region generation model.
[0190] The second extraction unit is used to extract the image corresponding to the target region in the adjusted image;
[0191] The training unit is used to input the extracted image into the hotspot region generation model for retraining;
[0192] The replacement unit is used to replace the original hotspot region generation model with the trained hotspot region generation model.
[0193] Please see Figure 6 , Figure 6 This is a schematic diagram of the structure of a hotspot annotation device 60 for a three-dimensional virtual sandbox provided in this application embodiment. The hotspot annotation device 60 for the three-dimensional virtual sandbox can be the aforementioned electronic device or a component within an electronic device. The hotspot annotation device 60 for the three-dimensional virtual sandbox includes: a processor 601, a communication interface 602, and a memory 603. The processor 601, communication interface 602, and memory 603 can be connected via a bus or other means; this application embodiment uses a bus connection as an example.
[0194] The processor 601 is the computing and control core of the hotspot marking device 60 in the 3D virtual sandbox. It can parse various instructions and data within the hotspot marking device 60. For example, the processor 601 can be a central processing unit (CPU), capable of transmitting various interactive data between internal structures within the hotspot marking device 60. The communication interface 602 can optionally include standard wired or wireless interfaces (such as Wi-Fi or mobile communication interfaces), and is controlled by the processor 601 to send and receive data. The communication interface 602 can also be used for the transmission and interaction of internal signaling or instructions within the hotspot marking device 60. The memory 603 is the storage device in the hotspot marking device 60, used to store programs and data. It is understood that the memory 603 here may include the built-in memory of the hotspot marking device 60 of the 3D virtual sandbox, or it may include the extended memory supported by the hotspot marking device 60 of the 3D virtual sandbox. The memory 603 provides storage space that stores the operating system of the hotspot marking device 60 of the 3D virtual sandbox. The storage space also stores the program code or instructions required by the processor to perform corresponding operations. Optionally, the storage space may also store relevant data generated by the processor after performing the corresponding operation.
[0195] In this embodiment, the processor 601 runs executable program code in the memory 603 to perform the following operations:
[0196] Extract one or more hotspot regions from the first frame image, wherein the first frame image is a frame image in the frame animation image sequence of the three-dimensional virtual sandbox in the web page;
[0197] The hotspot region generation model is used to retrieve target regions in each frame of the frame animation image sequence whose similarity to the hotspot region is higher than a preset threshold.
[0198] The hotspot region and the target region with a similarity higher than a preset threshold are marked as the same object;
[0199] The three-dimensional virtual sand table is displayed by playing the frame animation image sequence. During the display of the three-dimensional virtual sand table, the hot spots and target areas corresponding to the same object are highlighted using the same markings.
[0200] In one alternative embodiment, regarding the retrieval of target regions in each frame of the frame animation image sequence whose similarity to the hotspot region is higher than a preset threshold using a hotspot region generation model, the processor 601 is specifically configured to:
[0201] Extract the hotspot region to obtain an image corresponding to the hotspot region;
[0202] The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are input into the hotspot area generation model to obtain the initial target area corresponding to the hotspot area in each frame of the frame animation image sequence. The hotspot area generation model is trained based on the image corresponding to the historical hotspot area, the frame animation image sequence of the historical 3D virtual sandbox, and the target area corresponding to the historical hotspot area in each frame of the frame animation image sequence of the historical 3D virtual sandbox. The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are feature data, and the target area corresponding to the hotspot area in any frame of the frame animation image sequence is label data.
[0203] Initial target regions that have a similarity to the hotspot regions higher than a preset threshold are identified as target regions.
[0204] In one alternative embodiment, before inputting the image corresponding to the hotspot region and the frame animation image sequence of the 3D virtual sandbox into the hotspot region generation model to obtain the initial target region corresponding to the hotspot region in each frame of the frame animation image sequence, the processor 601 is further configured to:
[0205] Deep recognition is performed on images corresponding to historical hotspot areas and frame animation image sequences of historical 3D virtual sand table to determine the data information of the images corresponding to the historical hotspot areas and the frame animation image sequences of historical 3D virtual sand table, wherein the data information includes grayscale and resolution;
[0206] Based on the data information of the images corresponding to the historical hotspot areas and the frame animation image sequence of the historical 3D virtual sand table, adjust the data information of the images corresponding to the hotspot areas and the frame animation image sequence of the 3D virtual sand table.
[0207] In one alternative embodiment, the processor 601 is specifically configured to: determine the initial target region whose similarity to the hotspot region is higher than a preset threshold as the target region;
[0208] The frame of the animated image sequence is determined to include at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold.
[0209] The initial target regions with a similarity higher than a preset threshold to the hotspot region are sorted in descending order of their matching degree with the detailed information of the hotspot region, wherein the detailed information includes at least material information, shape information and color information;
[0210] Among the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold, the initial target region with the highest degree of matching with the detailed information of the hotspot region is selected as the target region.
[0211] In one alternative embodiment, regarding the sorting of the at least two initial target regions in descending order of similarity to the hotspot regions, the processor 601 is specifically configured to:
[0212] Obtain detailed information about the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than a preset threshold;
[0213] The detailed information matching degree between the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold is determined by comparing them with the detailed information of the hotspot region.
[0214] The initial target regions that are more than two times similar to the hotspot region are sorted in descending order of the degree of matching between the detailed information of the hotspot region and the detailed information of the hotspot region.
[0215] In one alternative embodiment, the processor 601 is further configured to:
[0216] The need to obtain training data for the hotspot region generation model;
[0217] The images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model are formatted to meet the training data requirements of the hotspot region generation model.
[0218] Extract the image corresponding to the target region from the adjusted image;
[0219] The extracted image is input into the hotspot region generation model for retraining;
[0220] Replace the original hotspot generation model with the trained hotspot generation model.
[0221] This application provides a computer-readable storage medium storing a computer program. The computer program includes program instructions, which, when executed by a processor, cause the processor to perform... Figure 2 and Figure 4 The operation performed.
[0222] This application also provides a computer program product that, when run on a processor, implements... Figure 2 and Figure 4 The operation performed.
[0223] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
Claims
1. A method for hotspot annotation in a three-dimensional virtual sandbox, characterized in that, The method is applied to a browser web page, and the method includes: Extract one or more hotspot regions from the first frame image, wherein the first frame image is a frame image in the frame animation image sequence of the three-dimensional virtual sandbox in the web page; Extract the hotspot region to obtain an image corresponding to the hotspot region; The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are input into the hotspot area generation model to obtain the initial target area corresponding to the hotspot area in each frame of the frame animation image sequence. The hotspot area generation model is trained based on the image corresponding to the historical hotspot area, the frame animation image sequence of the historical 3D virtual sandbox, and the target area corresponding to the historical hotspot area in each frame of the frame animation image sequence of the historical 3D virtual sandbox. The image corresponding to the hotspot area and the frame animation image sequence of the 3D virtual sandbox are feature data, and the target area corresponding to the hotspot area in any frame of the frame animation image sequence is label data. An initial target region whose similarity to the hotspot region is higher than a preset threshold is identified as the target region; the three-dimensional virtual sandbox is used for display in a browser webpage, and the hotspot region generation model is deployed in the browser webpage; The hotspot region and the target region with a similarity higher than a preset threshold are marked as the same object; The three-dimensional virtual sand table is displayed by playing the frame animation image sequence. During the display of the three-dimensional virtual sand table, the hot spots and target areas corresponding to the same object are highlighted using the same markings. The step of determining initial target regions with a similarity higher than a preset threshold to the hotspot regions as target regions includes: The frame of the animated image sequence is determined to include at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold. The initial target regions with a similarity higher than a preset threshold to the hotspot region are sorted in descending order of their matching degree with the detailed information of the hotspot region, wherein the detailed information includes at least material information, shape information and color information; Among the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold, the initial target region with the highest degree of matching with the detailed information of the hotspot region is selected as the target region.
2. The method according to claim 1, characterized in that, Before inputting the image corresponding to the hotspot region and the frame animation image sequence of the 3D virtual sandbox into the hotspot region generation model to obtain the initial target region corresponding to the hotspot region in each frame of the frame animation image sequence, the method further includes: Deep recognition is performed on images corresponding to historical hotspot areas and frame animation image sequences of historical 3D virtual sand table to determine the data information of the images corresponding to the historical hotspot areas and the frame animation image sequences of historical 3D virtual sand table, wherein the data information includes grayscale and resolution; Based on the data information of the images corresponding to the historical hotspot areas and the frame animation image sequence of the historical 3D virtual sand table, adjust the data information of the images corresponding to the hotspot areas and the frame animation image sequence of the 3D virtual sand table.
3. The method according to claim 1, characterized in that, The step of sorting the at least two initial target regions with a similarity higher than a preset threshold to the hotspot region according to their matching degree of detailed information with the hotspot region from high to low includes: Obtain detailed information about the at least two initial target regions and hotspot regions whose similarity to the hotspot region is higher than a preset threshold; The detailed information matching degree between the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold is determined by comparing them with the detailed information of the hotspot region. The initial target regions that are more than two times similar to the hotspot region are sorted in descending order of the degree of matching between the detailed information of the hotspot region and the detailed information of the hotspot region.
4. The method according to claim 1, characterized in that, After displaying the 3D virtual sandbox by playing the frame animation image sequence, and highlighting the hotspot areas and target areas corresponding to the same object using the same markers during the display of the 3D virtual sandbox, the method further includes: The need to obtain training data for the hotspot region generation model; The images in the frame animation image sequence corresponding to the target region output by the hotspot region generation model are formatted to meet the training data requirements of the hotspot region generation model. Extract the image corresponding to the target region from the adjusted image; The extracted image is input into the hotspot region generation model for retraining; Replace the original hotspot generation model with the trained hotspot generation model.
5. A hotspot marking device for a three-dimensional virtual sandbox, characterized in that, The device includes: The first extraction unit is used to extract one or more hot spots in the first frame image, wherein the first frame image is a frame image in the frame animation image sequence of the three-dimensional virtual sandbox in the web page; A retrieval unit is used to extract the hotspot region to obtain the image corresponding to the hotspot region; input the image corresponding to the hotspot region and the frame animation image sequence of the 3D virtual sandbox into the hotspot region generation model to obtain the initial target region corresponding to the hotspot region in each frame of the frame animation image sequence, wherein the hotspot region generation model is trained based on the image corresponding to the historical hotspot region, the frame animation image sequence of the historical 3D virtual sandbox, and the target region corresponding to the historical hotspot region in each frame of the frame animation image sequence of the historical 3D virtual sandbox, wherein the image corresponding to the hotspot region and the frame animation image sequence of the 3D virtual sandbox are feature data, and the target region corresponding to the hotspot region in any frame of the frame animation image sequence is label data; the initial target region with a similarity higher than a preset threshold to the hotspot region is determined as the target region; the 3D virtual sandbox is used to be displayed in a browser webpage, and the hotspot region generation model is deployed in the browser webpage; A marking unit is used to mark the hotspot region and a target region with a similarity to the hotspot region higher than a preset threshold as the same object; The playback unit is used to display the three-dimensional virtual sand table by playing the sequence of frame animation images. During the display of the three-dimensional virtual sand table, the same markings are used to highlight the hot spots and target areas corresponding to the same object. The retrieval unit is specifically used to determine that a frame of the frame animation image sequence includes at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold. The initial target regions with a similarity higher than a preset threshold to the hotspot region are sorted in descending order of their matching degree with the detailed information of the hotspot region, wherein the detailed information includes at least material information, shape information and color information; Among the at least two initial target regions whose similarity to the hotspot region is higher than a preset threshold, the initial target region with the highest degree of matching with the detailed information of the hotspot region is selected as the target region.
6. A hotspot marking device for a three-dimensional virtual sandbox, characterized in that, The hotspot marking device of the three-dimensional virtual sandbox includes at least one processor, a communication interface, and a memory. The communication interface is used to send and / or receive data, the memory is used to store computer programs, and the at least one processor is used to call at least one computer program stored in the memory to implement the method as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when run on a processor, implements the method as described in any one of claims 1-4.