Script scene recognition method and device, electronic equipment and readable storage medium
Patent Information
- Application Number
- CN202311368308.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-20
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-10-20
AI Technical Summary
但是剧本中场景数据通常较多,人工识别耗时耗力,效率较低
[0017]在本申请实施例中,对于剧本的多个场景数据中每一个场景数据,基于预先训练的识别模型对场景数据进行信息提取处理,得到场景信息,场景信息包括地点,或,场景信息包括地点和方位;在场景信息包括地点的情况下,将地点确定为场景数据对应的场景,在场景信息包括地点和方位的情况下,基于地点和方位确定场景数据对应的场景。通过上述方法,基于识别模型可以自动地提取场景数据中的场景信息,效率更高且准确度更高。基于场景信息即可确定场景数据对应的场景,提高了对剧本场景进行识别的效率和准确度。
Smart Images

Figure CN117521003B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of script scene technology, and in particular to a method, apparatus, electronic device, and readable storage medium for identifying script scenes. Background Technology
[0002] Each scene in the script includes scene data corresponding to the shooting location of that scene. Before filming a film or television drama, it is usually necessary to identify the scene data of the script to determine the main and secondary scenes corresponding to the shooting of the drama, so as to arrange the shooting locations in advance and calculate the shooting costs.
[0003] Currently, the process typically involves manually identifying scene data in the script to obtain the locations and orientations described within the scene data, and then combining these locations and orientations to create the main and secondary shooting scenes. However, scripts usually contain a large amount of scene data, making manual identification time-consuming, labor-intensive, and inefficient. Summary of the Invention
[0004] The purpose of this invention is to provide a method, apparatus, electronic device, and readable storage medium for identifying script scenes, thereby improving the efficiency of script scene identification. The specific technical solution is as follows: In a first aspect of this invention, a method for identifying script scenes is provided, comprising: For each scene data in the script's multiple scene data, information extraction processing is performed on the scene data based on a pre-trained recognition model to obtain scene information. The scene information includes location, or the scene information includes location and orientation, where the orientation is used to characterize the orientation information of the shooting location corresponding to the scene data at the location. If the scene information includes a location, the location is determined as the scene corresponding to the scene data. If the scene information includes both location and orientation, the scene corresponding to the scene data is determined based on the location and orientation.
[0005] Optionally, the method further includes: When there is only one scene, the scene is determined as the main scene corresponding to the scene data; when there are at least two scenes, the main scene and the secondary scene corresponding to the scene data are determined based on the at least two scenes.
[0006] Optionally, the recognition model includes a location recognition sub-model, and the scene information is extracted and processed based on the pre-trained recognition model to obtain scene information, including: The scene data is input into the location identification sub-model for information annotation processing to obtain first annotation information, or the first annotation information and second annotation information. The first annotation information is used to indicate the location in the scene data, and the second annotation information is used to indicate the orientation in the scene data. The location is extracted from the scene data based on the indication of the first annotation information, or the location and orientation are extracted from the scene data based on the indication of the first annotation information and the second annotation information.
[0007] Optionally, the method further includes: Obtain multiple first training scenarios, each of which carries location annotation information and / or orientation annotation information; Based on the multiple first training scenarios, the location identification sub-model is iteratively trained to obtain the location identification sub-model.
[0008] Optionally, the recognition model further includes a regional recognition sub-model, the scene information further includes regional information, and the step of extracting information from the scene data based on the pre-trained recognition model to obtain scene information further includes: The scene data is input into the region identification sub-model for information annotation processing to obtain third annotation information, which is used to indicate the region in the scene data. Based on the indication of the third annotation information, the regional information is extracted from the scene data.
[0009] Optionally, the method further includes: Multiple second training scenarios are acquired, each carrying geographic labeling information; Based on the multiple second training scenarios, the region identification sub-model to be trained is iteratively trained to obtain the region identification sub-model.
[0010] Optionally, the step of inputting the scene data into the location identification sub-model for information annotation processing to obtain first annotation information, or obtaining the first annotation information and second annotation information, includes: The geographic information in the scene data is separated to obtain the target data; The target data is input into the location identification sub-model for information annotation processing to obtain first annotation information, or the first annotation information and second annotation information.
[0011] Optionally, after extracting the geographic information from the scene data based on the indication of the third annotation information, the method further includes: Obtain the target ratio corresponding to the regional information. The target ratio is the proportion of the number of second scene data to the number of first scene data. The first scene data is the scene data containing the region among multiple scene data. The second scene data is the first scene data in which the regional information is successfully extracted. If the target ratio is greater than or equal to the threshold, the regional information is extracted from the third scene data; if the target ratio is less than the threshold, the regional information extracted from the second scene data is deleted. The third scene data is the first scene data in which the regional information was not successfully extracted.
[0012] Optionally, when the number of scenes is one, the scene is determined as the main scene corresponding to the scene data; when the number of scenes is at least two, after determining the main scene and secondary scene corresponding to the scene data based on the at least two scenes, the method further includes: The target scene is obtained by merging and matching the main scenes corresponding to multiple scene data. If the main scene of the scene data includes the target location, the target scene in the main scene is determined as the target main scene, and the content in the main scene other than the target scene is determined as the target secondary scene.
[0013] Optionally, the step of merging and matching the main scenes corresponding to multiple scene data to obtain the target scene includes: For each main scene, determine the number of main scenes that match the main scene among the multiple main scenes corresponding to the scene data; The main scene with the most matching main scenes is identified as the target scene.
[0014] In a second aspect of the invention, a script scene recognition device is provided, comprising: The first processing module is used to extract information from each scene data in the script based on a pre-trained recognition model to obtain scene information. The scene information includes location, or the scene information includes location and orientation, and the orientation is used to characterize the orientation information of the shooting position corresponding to the scene data at the location. The first determining module is used to determine the location as the scene corresponding to the scene data when the scene information includes a location, and to determine the scene corresponding to the scene data based on the location and the orientation when the scene information includes a location and an orientation.
[0015] In a third aspect of the present invention, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus. Memory, used to store programs; When a processor executes a program stored in memory, it implements the steps of the method as described in the first aspect.
[0016] In a fourth aspect of the invention, a computer-readable storage medium is provided having a program stored thereon that, when executed by a processor, implements the steps of the method as described in the first aspect.
[0017] In this embodiment, for each scene data point in the script's multiple scene data sets, information extraction processing is performed based on a pre-trained recognition model to obtain scene information. The scene information includes location, or it includes both location and orientation. If the scene information includes location, the location is identified as the scene corresponding to the scene data; if the scene information includes both location and orientation, the scene corresponding to the scene data is identified based on both location and orientation. Through this method, scene information can be automatically extracted from the scene data based on the recognition model, resulting in higher efficiency and accuracy. The scene corresponding to the scene data can be determined based on the scene information, improving the efficiency and accuracy of script scene recognition. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0019] Figure 1 This is one of the flowcharts illustrating the script scene recognition method in an embodiment of the present invention; Figure 2 This is a schematic diagram of the recognition model training process in an embodiment of the present invention; Figure 3 This is a second flowchart illustrating the script scene recognition method in an embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of the script scene recognition device in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of these steps can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the steps can be rearranged. A process can be terminated when its operation is complete, but it may also have additional steps not included in the figures. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.
[0022] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0023] In practical applications, the script is divided according to shooting scenes. The script content corresponding to each shooting scene contains descriptive content of the shooting scene in which the shooting scene takes place. The scene data is the content in the script used to describe the specific shooting location of the scene.
[0024] Please see Figure 1 , Figure 1 This is one of the flowcharts illustrating the script scene identification method in this embodiment of the invention. The script scene identification method provided in this embodiment of the invention can be used to identify shooting scenes contained in different shooting scenes in a script, thereby obtaining the scene corresponding to each shooting scene in the script.
[0025] like Figure 1 As shown, the method for recognizing script scenes specifically includes the following steps: Step 101: For each scene data in the script's multiple scene data, perform information extraction processing on the scene data based on a pre-trained recognition model to obtain scene information. The scene information includes location, or the scene information includes location and orientation, where the orientation is used to characterize the orientation information of the shooting location corresponding to the scene data at the location.
[0026] Step 102: If the scene information includes a location, determine the location as the scene corresponding to the scene data; if the scene information includes a location and a direction, determine the scene corresponding to the scene data based on the location and the direction.
[0027] For each scene data point in the script, an information extraction process is performed based on a pre-trained recognition model to obtain the scene information contained in each scene data point. This scene information may include only the location, or it may include both the location and its orientation.
[0028] It should be understood that "location" is used to characterize the location of the shooting scene, usually a space or a small area or location, such as "office," "restaurant," "street," and "room." Since a location is usually a space or scene, when the space or area is large, the shooting location may need to be refined to a specific direction within that location. In this case, the direction can be used to characterize the shooting location's orientation within the location, such as "inside," "outside," "side," and "doorway."
[0029] In practical implementation, when scene information includes location and orientation, orientation usually appears in conjunction with location. However, the specific content of orientation can vary depending on the location. For example, "inside a room," "at the street corner," and "beside the gate of a house."
[0030] To make it easier to understand, examples will be given below.
[0031] For example, the scene data includes the content "corridor outside the office". Based on the recognition model, the scene data can be processed to extract information and obtain two locations, "office" and "corridor", as well as a direction "outside". Since "outside" follows "office", this direction is used to characterize the specific location of the shooting position within the office.
[0032] For example, the scene data includes the content "in front of the ticket window of the train station". Based on the recognition model, the scene data can be processed to extract information to obtain a location "ticket window of the train station" (or it may be possible to extract two locations, "train station" and "ticket window"), and a direction "in front". Since "in front" follows "ticket window", this direction is used to represent the specific location of the shooting position in front of the ticket window.
[0033] In some cases, scene information only includes the location. In such cases, the location is directly identified as the scene corresponding to the scene data. It should be noted that the number of locations is usually one, but depending on the script, the number of locations can be multiple; this is not limited here.
[0034] For example, in some embodiments, the scene information includes "office," then the scene corresponding to the scene data is "office." In other embodiments, the scene information includes "office building" and "company," then the scene corresponding to the scene data is "office building" and "company."
[0035] In other cases, scene information includes location and orientation. In these situations, the scene corresponding to the scene data is determined based on the location and orientation. It should be noted that the number of locations is not limited here. When there are multiple locations, orientation can also be used to represent the association between two addresses.
[0036] For example, in some embodiments, the scene information includes the location "office" and the orientation "inside". In this case, the scene corresponding to the scene data can be determined as "inside the office". In other embodiments, the scene information includes the location "office" and "tea room", as well as the orientation "inside" corresponding to "office". In this case, the scene corresponding to the scene data can be determined as "inside the office" and "tea room", or the scene corresponding to the scene data can be determined as "tea room inside the office". The specific settings can be made according to actual business needs, which will not be elaborated here.
[0037] In this embodiment, for each scene data in the script's multiple scene data sets, information extraction processing is performed on the scene data based on a pre-trained recognition model to obtain scene information. The scene information includes location, or it includes both location and orientation. If the scene information includes location, the location is identified as the scene corresponding to the scene data; if the scene information includes both location and orientation, the scene corresponding to the scene data is identified based on both location and orientation. This method automatically extracts scene information from the scene data based on the recognition model, resulting in higher efficiency and accuracy. The scene corresponding to the scene data can be determined based on the scene information, improving the efficiency and accuracy of script scene recognition.
[0038] Optionally, in some embodiments, the method further includes: when the number of scenes is one, determining the scene as the main scene corresponding to the scene data; when the number of scenes is at least two, determining the main scene and secondary scene corresponding to the scene data based on the at least two scenes.
[0039] As mentioned above, depending on the actual business needs, the number of scenarios corresponding to each data scenario can be one, two, or more. To facilitate scenario management, when there are multiple scenarios, they are divided into one main scenario and multiple sub-scenarios. As one possible division method, the main scenario is typically a large, overall scenario, and the sub-scenarios are parts of the main scenario. Another possible division method is to define the main scenario as the scenario that appears most frequently.
[0040] When there is only one scene, that scene is the primary scene corresponding to the scene data. When there are at least two scenes, one of the at least two scenes is designated as the primary scene, and the other scenes are designated as secondary scenes. The specific method for determining the primary and secondary scenes corresponding to the scene data based on at least two scenes is not limited here.
[0041] For example, as an optional implementation, a primary scene and a secondary scene can be determined based on the association between at least two scenes. For example, the scene data includes "tea room in the office," and the scenes corresponding to this scene data are "inside the office" and "tea room." Since the tea room is located inside the office, "inside the office" can be determined as the primary scene, and "tea room" can be determined as the secondary scene.
[0042] As an alternative implementation, primary and secondary scenes can be determined based on the frequency of each scene's appearance in the script. It should be understood that different scenes from the same script may be filmed in the same location, therefore different scene data may correspond to the same scene. For example, if the scene "inside the office" appears 50 times in all scene data of the script, and "in the break room" appears 5 times, then "inside the office" is determined as the primary scene, and "in the break room" as the secondary scene.
[0043] In this embodiment, when there is only one scene, the scene is designated as the primary scene corresponding to the scene data. When there are at least two scenes, the primary scene and secondary scene corresponding to the scene data are determined based on at least two scenes. This method facilitates unified management of the scenes corresponding to the scene data.
[0044] Optionally, when the number of scenes is one, the scene is determined as the main scene corresponding to the scene data; when the number of scenes is at least two, after determining the main scene and secondary scene corresponding to the scene data based on the at least two scenes, the method further includes: The target scene is obtained by merging and matching the main scenes corresponding to multiple scene data. If the main scene of the scene data includes the target location, the target scene in the main scene is determined as the target main scene, and the content in the main scene other than the target scene is determined as the target secondary scene.
[0045] In practice, due to factors such as the effectiveness of the recognition model, the difficulty of recognizing scene information, and actual business needs, different scenes may be obtained for the same scene data when the information is extracted and processed based on the pre-trained recognition model.
[0046] To facilitate understanding, examples will be provided below. For instance, some scene data all contain the content "tea room in the office", some scene data all contain the content "meeting room in the office", and some scene data all contain the content "in the office". Information extraction and processing are performed on these scene data based on the pre-trained recognition model.
[0047] As mentioned above, for scene data containing "the tea room in the office," depending on the actual processing method, the main scene corresponding to this part of the scene data may be "inside the office," "the tea room," or "the tea room in the office." Similarly, for scene data containing "the meeting room in the office," depending on the actual processing method, the main scene corresponding to this part of the scene data may be "inside the office," "the meeting room," or "the meeting room in the office." And for scene data containing "inside the office," the main scene is "inside the office."
[0048] For the same script, unifying the main scene for different scene data makes it easier to manage and arrange the main scenes for different scene data, and also makes it easier to calculate the shooting costs in advance. Therefore, in the example above, we want the main scene for scene data containing the content "inside the office" to be "inside the office".
[0049] As a specific embodiment, the shortest prefix matching algorithm is used to merge and match the main scenes corresponding to multiple scene data to obtain the target scene. Optionally, in some embodiments, merging and matching the main scenes corresponding to multiple scene data to obtain the target scene includes: For each main scene, determine the number of main scenes that match the main scene among the multiple main scenes corresponding to the scene data; The main scene with the most matching main scenes is identified as the target scene.
[0050] For example, when the main scenarios of multiple scenario data include the categories of "meeting room in the office", "tea room in the office" and "in the office", the number of main scenarios matching "meeting room in the office", "tea room in the office" and "in the office" are determined in turn.
[0051] For the main scene "Inside the Office", since the three main scenes "Meeting Room in the Office", "Tea Room in the Office" and "Inside the Office" all contain "Inside the Office", the number of main scenes that match the main scene "Inside the Office" is 3. Similarly, for the main scenes "Meeting Room in the Office" and "Tea Room in the Office", the number of main scenes that match is 1. Therefore, the target scene is "Inside the Office".
[0052] Since the current main scene of the scene data, "Meeting Room in the Office," includes "Inside the Office," the original main scene of this scene data is split into the target main scene "Inside the Office" and the target secondary scene "Meeting Room." Similarly, since "Coffee Room in the Office" includes "Inside the Office," it is determined as the target main scene "Inside the Office" and the target secondary scene "Coffee Room." For "Inside the Office," since it does not contain any other content besides the target scene, its corresponding target main scene remains "Inside the Office" and does not include the target secondary scene.
[0053] In this embodiment, the existing main scene is used as the shortest character for merging, which avoids forcibly splitting words in the main scene and affecting its accuracy. For ease of understanding, an example is given below. Suppose that the main scenes in multiple scene data include "meeting room in the office," "tea room in the office," and "lounge room in the office." Each of these three main scenes only matches itself, i.e., the quantity is 1. In this case, it is impossible to merge these three main scenes.
[0054] Of course, in some embodiments, the location can also be used as the shortest character. For each location, the number of main scenes that match the location in the main scenes corresponding to the multiple scene data is determined; the location with the most matching main scenes is determined as the target scene.
[0055] In this embodiment, the main scenes corresponding to multiple scene data are merged and matched to obtain the target scene. If the main scene of the scene data includes the target scene, the main scene of the scene data is split into the target main scene and the target secondary scene. This method unifies the main scenes from multiple scene data sets, facilitating the management of different scenes corresponding to the same script and improving the accuracy of script scene recognition.
[0056] Optionally, in some embodiments, the recognition model includes a location identification sub-model, and the scene information is extracted and processed based on the pre-trained recognition model to obtain scene information, including: The scene data is input into the location identification sub-model for information annotation processing to obtain first annotation information, or the first annotation information and second annotation information. The first annotation information is used to indicate the location in the scene data, and the second annotation information is used to indicate the orientation in the scene data. The location is extracted from the scene data based on the indication of the first annotation information, or the location and orientation are extracted from the scene data based on the indication of the first annotation information and the second annotation information.
[0057] It should be understood that the first annotation information is used to mark the specific location of the words corresponding to the location in the scene data, and the second annotation information is used to mark the specific location of the words corresponding to the direction in the scene data. Based on the first annotation information and the second annotation information, the location and direction can be extracted from the corresponding locations in the scene data.
[0058] In some embodiments, the location identification sub-model is a pre-trained deep learning model. Optionally, in some embodiments, the method further includes: Obtain multiple first training scenarios, each of which carries location annotation information and / or orientation annotation information; Based on the multiple first training scenarios, the location identification sub-model is iteratively trained to obtain the location identification sub-model.
[0059] In this embodiment, a script with regular scene descriptions can be obtained, and the scenes in the script can be annotated by a program or manually to obtain a first training scene. The first training scene has been pre-annotated with location annotation information and / or orientation annotation information.
[0060] For example, the location identification sub-model is a Transformer, which labels the locations and orientations in scene data with fixed rules to generate training data, and then uses this part of labeled training data to train the Transformer to obtain the location identification sub-model.
[0061] In one scenario, the scene data only includes locations. The location identification sub-model annotates the locations in the scene data to obtain first annotation information, and extracts the location from the corresponding location in the scene data based on the indication of the first annotation information.
[0062] In another scenario, the scene data includes location and orientation. The location and orientation recognition sub-model labels the location and orientation in the scene data to obtain first labeling information and second labeling information. Based on the first labeling information, the location is extracted from the corresponding position in the scene data, and based on the second labeling information, the orientation is extracted from the corresponding position in the scene data.
[0063] To facilitate understanding, an example will be provided below. For instance, the scene data includes: "This scene was filmed in the corridor outside the office". The scene data is input into the location recognition sub-model for information annotation processing to obtain first annotation information and second annotation information. The first annotation information is used to indicate the location of "office" and "corridor" in the scene data, and the second annotation information is used to indicate the location of "outside" in the scene data. Based on the first annotation information and the second annotation information, "office", "outside", and "corridor" can be extracted.
[0064] In this embodiment, the above method is used to perform information annotation processing on scene data using a location identification sub-model, thereby extracting scene information from the scene data and improving the efficiency and accuracy of extracting scene information from scene data.
[0065] Optionally, in some embodiments, the recognition model further includes a regional recognition sub-model, the scene information further includes regional information, and the step of performing information extraction processing on the scene data based on the pre-trained recognition model to obtain scene information further includes: The scene data is input into the region identification sub-model for information annotation processing to obtain third annotation information, which is used to indicate the region in the scene data. Based on the indication of the third annotation information, the regional information is extracted from the scene data.
[0066] It should be understood that the third annotation information is used to mark the specific location of the words corresponding to the region in the scene data. Based on the third annotation information, the region can be extracted from the corresponding location in the scene data.
[0067] In practical implementation, the filming of certain movies and TV series may require shooting in different locations. For example, even if both locations are train stations, "Nanjing Railway Station" and "Beijing Railway Station" should actually be two different filming scenes. In this embodiment, the scene information also includes geographical information, that is, the scene information includes geographical information and location, or the scene information includes geographical information, location, and orientation.
[0068] In some embodiments, the region identification sub-model is a pre-trained deep learning model. Optionally, in some embodiments, the method further includes: Multiple second training scenarios are acquired, each carrying geographic labeling information; Based on the multiple second training scenarios, the region identification sub-model to be trained is iteratively trained to obtain the region identification sub-model.
[0069] In this embodiment, a script with regular scene descriptions can be obtained, and the scenes in the script can be annotated by a program or manually to obtain a second training scene. The second training scene has been pre-annotated with regional annotation information.
[0070] For example, the region recognition sub-model is a Transformer. Regions in scene data with fixed rules are labeled to generate training data. Then, the Transformer is trained using this labeled training data to obtain the region recognition sub-model.
[0071] The region identification sub-model can annotate the region information in the scene data to obtain third annotation information. For example, the scene data includes: "This shooting took place at the entrance of the Sichuan Provincial Museum". The scene data is input into the region identification sub-model for information annotation processing to obtain third annotation information. The third annotation information is used to mark the location of "Sichuan Province" in the scene data. Based on the third annotation information, "Sichuan Province" can be extracted.
[0072] In this embodiment, the recognition model further includes a regional recognition sub-model, and the scene information also includes regional information. Through the above method, regional information can be further extracted from the scene data, making the scene division more accurate. Simultaneously, using the regional recognition sub-model to extract regional information from the scene data improves the efficiency and accuracy of extracting regional information from the scene data.
[0073] As an optional implementation method, scene data is input into the location identification sub-model and the region identification sub-model respectively. The location identification sub-model is used to extract the location and orientation in the data scene, and the region identification sub-model is used to extract the region information in the data scene, thereby improving the efficiency of obtaining scene information.
[0074] Optionally, as another optional implementation, the step of inputting the scene data into the location identification sub-model for information annotation processing to obtain first annotation information, or obtaining the first annotation information and second annotation information, includes: The geographic information in the scene data is separated to obtain the target data; The target data is input into the location identification sub-model for information annotation processing to obtain first annotation information, or the first annotation information and second annotation information.
[0075] In practice, the words corresponding to geographical information and location may be easily confused, causing the location identification sub-model to mistakenly label geographical information as location when performing information labeling, thus affecting the accuracy of scene information extraction.
[0076] To facilitate understanding, an example will be provided below. For instance, the scene data includes: "This filming took place at the entrance of the Sichuan Provincial Museum." The scene data is input into the regional identification sub-model for information annotation processing, extracting the regional information "Sichuan Province." "Sichuan Province" is then removed from the scene data, resulting in the target data "This filming took place at the entrance of the museum." "This filming took place at the entrance of the museum" is then input into the location identification sub-model for information annotation processing, obtaining first annotation information and second annotation information. Based on the first annotation information, the location "Museum" is extracted, and based on the second annotation information, the orientation "Entrance" is extracted.
[0077] In this embodiment, after extracting regional information from the scene data based on the indication of the third annotation information, the regional information in the scene data is separated and deleted to obtain the target data. The target data is then input into the location identification sub-model for information annotation processing to obtain first annotation information, or the first annotation information and second annotation information. This method reduces the impact of regional information on the information annotation processing of the location identification sub-model, improves the accuracy of the extracted locations, and thus improves the accuracy of scene information extraction.
[0078] It should be understood that in some non-realistic or historical drama scripts, geographical information may be fictional. For example, in fantasy dramas, geographical information might be "Qingqiu City," "Country M," or "City A." For this type of fictional geographical information, the geographical identification sub-model may have issues with incomplete or incorrect extraction; that is, the geographical information may not be correctly extracted from some scene data, or non-geographical information from some scene data may be extracted as geographical information.
[0079] Optionally, after extracting the geographic information from the scene data based on the indication of the third annotation information, the method further includes: Obtain the target ratio corresponding to the regional information. The target ratio is the proportion of the number of second scene data to the number of first scene data. The first scene data is the scene data containing the region among multiple scene data. The second scene data is the first scene data in which the regional information is successfully extracted. If the target ratio is greater than or equal to the threshold, the regional information is extracted from the third scene data; if the target ratio is less than the threshold, the regional information extracted from the second scene data is deleted. The third scene data is the first scene data in which the regional information was not successfully extracted.
[0080] For any given geographic information, determine the number of first-scene data and second-scene data corresponding to that geographic information, and then obtain the target ratio. The target ratio can be used to characterize the probability that the geographic information is correctly identified and labeled by the geographic identification sub-model, and thus measure the probability that the geographic information is correct.
[0081] If the target ratio is greater than or equal to the threshold, the regional information can be considered to be likely correct and needs to be extracted. Therefore, this extracted regional information is used as the regional information corresponding to the third scenario data. Through the above processing, it is equivalent to extracting this regional information from all first scenario data, thus ensuring that the regional information is completely extracted.
[0082] If the target ratio is less than the threshold, the regional information is likely incorrect and should not be extracted. Therefore, the regional information already extracted from the second scene data is returned (equivalent to deleting the regional information extracted from the second scene data). Through the above processing, it is equivalent to not extracting the regional information from all first scene data, thus avoiding the extraction of erroneous regional information.
[0083] To facilitate understanding, an example will be provided below. For instance, the threshold is set to 0.5, and a script includes 100 scene data points, 50 of which contain "Country M". These 50 scene data points are then input into the regional identification sub-model for information annotation processing.
[0084] In one scenario, "Country M" in 40 scene data points was correctly labeled as geographical information and successfully extracted; the geographical information corresponding to these 40 scene data points was "Country M". "Country M" in the remaining 10 scene data points was not correctly labeled, and no geographical information was corresponding to these 10 scenarios. In this case, the target ratio is 4 / 5. Since the target ratio is greater than 0.5, "Country M" in the remaining 10 scene data points is extracted as geographical information, ensuring that all 50 scene data points correspond to the geographical information "Country M".
[0085] In another scenario, "Country M" in 20 scene data points was correctly labeled as geographical information and successfully extracted, while "Country M" in the remaining 30 scene data points was not correctly labeled. In this case, the target ratio is 2 / 5. Since the target ratio is less than 0.5, the geographical information corresponding to the 20 scene data points where "Country M" was successfully extracted is deleted, so that none of the 50 scene data points correspond to the geographical information "Country M".
[0086] In this embodiment, for each piece of geographic information, its validity is determined by the probability that it is correctly identified and extracted by the geographic identification sub-model in different scene data. This method avoids the problems of incomplete geographic information extraction and errors caused by fictitious geographic locations, improving the accuracy of geographic information extraction and optimizing the recognition effect of script scenes.
[0087] To facilitate understanding, a specific embodiment will be used below to illustrate the script scene recognition method provided by this invention. First, two Transformer models are pre-trained. Please refer to... Figure 2 Specifically, the program annotates geographic information in scenes with fixed rules, generating labeled data. This labeled data is then used to train a Transformer model, resulting in a Transformer recognition model for identifying geographic locations, i.e., a geographic identification sub-model. Conversely, the program annotates locations and orientations in scenes with fixed rules, generating labeled data. This labeled data is then used to train a Transformer model, resulting in a Transformer model for identifying locations and orientations, i.e., a location and orientation recognition sub-model.
[0088] Please see Figure 3 The specific steps for script scene recognition based on the pre-trained region recognition sub-model and location recognition sub-model are as follows: Step 1: Extract scene data from the script. A script typically includes multiple shooting scenes, corresponding to multiple scene data sets, such as... Figure 3 The scene data shown is 1, 2, ..., N, where N is a positive integer greater than 3.
[0089] Step 2: For each scene data, use the pre-trained regional recognition sub-model to perform information annotation processing on the scene data, annotating the regional information in the scene data to obtain scene data containing third annotation information.
[0090] Step 3: Identify the third annotation information and extract the regional information from the scene data.
[0091] Step 4: Address the issue of incomplete region extraction. The specific method is as follows: For all scene data, count the scene data within the same script that contain the same region information. If the number of scene data with extracted region information is greater than the number of scene data without extracted region information, then that region information is treated as the region information of the scene data without extracted region information. If the number of scene data with extracted region information is less than the number of scene data without extracted region information, then the extracted region information is placed back into the scene data where it was extracted, and that scene data does not correspond to that region information. After correcting the region information of all scene data according to Step 4, delete the region information corresponding to the scene data to obtain the target data after removing the region information.
[0092] Step 5: For the target data after removing the region, use the pre-trained location recognition sub-model to perform information annotation processing, annotate the location and orientation in the scene data (if the scene data contains orientation), and obtain scene data containing the first annotation information and the second annotation information (if the scene data contains orientation).
[0093] Step 6: Identify the first and second annotation information (if the scene data contains the orientation), and extract the location and orientation from the target data (if the scene data contains the orientation).
[0094] Step 7: Based on the business rules in actual application, determine the primary and secondary scenarios corresponding to each scenario data based on location and orientation (when the number of scenarios is greater than 1).
[0095] Step 8: Merge and match the main scenes corresponding to multiple scene data sets to obtain the target scene. Once the target scene is obtained, process the main and secondary scenes corresponding to the multiple scene data sets to achieve a unified main scene.
[0096] Step 9: Output the main scene, secondary scene (optional), and geographic information (optional) corresponding to each scene data.
[0097] In this embodiment, a recognition model is used to automatically identify geographical information, locations, and orientations within the script's scenes, improving accuracy and efficiency while reducing manual workload and errors, thereby lowering costs and increasing efficiency. Simultaneously, for fictitious geographical information, the extraction results are statistically analyzed to optimize the extraction process, improving accuracy. To address the issue of low consistency in main scenes caused by fictitious locations, a shortest prefix matching algorithm is used to match and merge the main scenes from multiple scene data sets, improving the consistency of the main scenes corresponding to different scene data sets within the same script.
[0098] like Figure 4As shown, this embodiment of the invention also provides a script scene recognition device 400, comprising: The first processing module 401 is used to perform information extraction processing on each scene data in the script based on a pre-trained recognition model to obtain scene information. The scene information includes location, or the scene information includes location and orientation. The orientation is used to characterize the orientation information of the shooting position corresponding to the scene data at the location. The first determining module 402 is used to determine the location as the scene corresponding to the scene data when the scene information includes a location, and to determine the scene corresponding to the scene data based on the location and the orientation when the scene information includes a location and an orientation.
[0099] Optionally, the script scene recognition device 400 further includes: The second determining module is used to determine the scene as the main scene corresponding to the scene data when the number of scenes is one; and to determine the main scene and secondary scene corresponding to the scene data based on the at least two scenes when the number of scenes is at least two.
[0100] Optionally, the recognition model includes a location recognition sub-model, and the first processing module 401 includes: The first processing unit is used to input the scene data into the location identification sub-model for information annotation processing to obtain first annotation information, or to obtain the first annotation information and second annotation information. The first annotation information is used to indicate the location in the scene data, and the second annotation information is used to indicate the orientation in the scene data. The first extraction unit is configured to extract the location from the scene data based on the indication of the first annotation information, or to extract the location and orientation from the scene data based on the indication of the first annotation information and the second annotation information.
[0101] Optionally, the script scene recognition device 400 further includes: The first acquisition module is used to acquire multiple first training scenarios, wherein the first training scenarios carry location annotation information and / or orientation annotation information; The first training module is used to iteratively train the location identification sub-model to be trained based on the multiple first training scenarios, so as to obtain the location identification sub-model.
[0102] Optionally, the recognition model further includes a regional recognition sub-model, the scene information further includes regional information, and the first processing module 401 further includes: The second processing unit is used to input the scene data into the region identification sub-model for information annotation processing to obtain third annotation information, which is used to indicate the region in the scene data. The second extraction unit is used to extract the regional information from the scene data based on the indication of the third annotation information.
[0103] Optionally, the script scene recognition device 400 further includes: The second acquisition module is used to acquire multiple second training scenarios, each carrying geographic labeling information. The second training module is used to iteratively train the region identification sub-model to be trained based on the multiple second training scenarios to obtain the region identification sub-model.
[0104] Optionally, the first extraction unit is specifically used for: The geographic information in the scene data is separated to obtain the target data; The target data is input into the location identification sub-model for information annotation processing to obtain first annotation information, or the first annotation information and second annotation information.
[0105] Optionally, the first processing module 401 further includes: An acquisition unit is used to acquire a target ratio corresponding to the regional information. The target ratio is the proportion of the number of second scene data to the number of first scene data. The first scene data is scene data containing the region among multiple scene data. The second scene data is the first scene data in which the regional information has been successfully extracted. The third processing unit is used to extract the regional information from the third scene data when the target ratio is greater than or equal to the threshold, and to delete the regional information extracted from the second scene data when the target ratio is less than the threshold. The third scene data is the first scene data in which the regional information was not successfully extracted.
[0106] Optionally, the script scene recognition device 400 further includes: The second processing module is used to merge and match the main scenes corresponding to multiple scene data to obtain the target scene. The third determining module is used to determine the target scene in the main scene as the target main scene when the main scene of the scene data includes the target location, and to determine the content in the main scene other than the target scene as the target secondary scene.
[0107] Optionally, the second processing module includes: The first determining unit is used to determine, for each main scene, the number of main scenes that match the main scene among the main scenes corresponding to the multiple scene data; The second determining unit is used to determine the main scene with the most matching main scenes as the target scene.
[0108] The script scene recognition device 400 provided in this application embodiment can implement the various processes implemented in the above method embodiment, and will not be described again here to avoid repetition.
[0109] This invention also provides an electronic device, such as... Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503, and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504. Memory 503 is used to store computer programs; When processor 501 executes the program stored in memory 503, it performs the following steps: For each scene data in the script's multiple scene data, information extraction processing is performed on the scene data based on a pre-trained recognition model to obtain scene information. The scene information includes location, or the scene information includes location and orientation, where the orientation is used to characterize the orientation information of the shooting location corresponding to the scene data at the location. If the scene information includes a location, the location is determined as the scene corresponding to the scene data. If the scene information includes both location and orientation, the scene corresponding to the scene data is determined based on the location and orientation.
[0110] Optionally, when executing a program stored in memory 503, processor 501 may also perform the following steps: When there is only one scene, the scene is determined as the main scene corresponding to the scene data; when there are at least two scenes, the main scene and the secondary scene corresponding to the scene data are determined based on the at least two scenes.
[0111] Optionally, the recognition model includes a location recognition sub-model. The processor 501 is also used to execute the program stored in the memory 503 to perform the following steps: The scene data is input into the location identification sub-model for information annotation processing to obtain first annotation information, or the first annotation information and second annotation information. The first annotation information is used to indicate the location in the scene data, and the second annotation information is used to indicate the orientation in the scene data. The location is extracted from the scene data based on the indication of the first annotation information, or the location and orientation are extracted from the scene data based on the indication of the first annotation information and the second annotation information.
[0112] Optionally, when executing a program stored in memory 503, processor 501 may also perform the following steps: Obtain multiple first training scenarios, each of which carries location annotation information and / or orientation annotation information; Based on the multiple first training scenarios, the location identification sub-model is iteratively trained to obtain the location identification sub-model.
[0113] Optionally, the recognition model further includes a regional recognition sub-model, and the scene information further includes regional information. The processor 501, when executing the program stored in the memory 503, performs the following steps: The scene data is input into the region identification sub-model for information annotation processing to obtain third annotation information, which is used to indicate the region in the scene data. Based on the indication of the third annotation information, the regional information is extracted from the scene data.
[0114] Optionally, when executing a program stored in memory 503, processor 501 may also perform the following steps: Multiple second training scenarios are acquired, each carrying geographic labeling information; Based on the multiple second training scenarios, the region identification sub-model to be trained is iteratively trained to obtain the region identification sub-model.
[0115] Optionally, when executing a program stored in memory 503, processor 501 may also perform the following steps: The geographic information in the scene data is separated to obtain the target data; The target data is input into the location identification sub-model for information annotation processing to obtain first annotation information, or the first annotation information and second annotation information.
[0116] Optionally, when executing a program stored in memory 503, processor 501 may also perform the following steps: Obtain the target ratio corresponding to the regional information. The target ratio is the proportion of the number of second scene data to the number of first scene data. The first scene data is the scene data containing the region among multiple scene data. The second scene data is the first scene data in which the regional information is successfully extracted. If the target ratio is greater than or equal to the threshold, the regional information is extracted from the third scene data; if the target ratio is less than the threshold, the regional information extracted from the second scene data is deleted. The third scene data is the first scene data in which the regional information was not successfully extracted.
[0117] Optionally, when executing a program stored in memory 503, processor 501 may also perform the following steps: The target scene is obtained by merging and matching the main scenes corresponding to multiple scene data. If the main scene of the scene data includes the target location, the target scene in the main scene is determined as the target main scene, and the content in the main scene other than the target scene is determined as the target secondary scene.
[0118] Optionally, when executing a program stored in memory 503, processor 501 may also perform the following steps: For each main scene, determine the number of main scenes that match the main scene among the multiple main scenes corresponding to the scene data; The main scene with the most matching main scenes is identified as the target scene.
[0119] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.
[0120] The communication interface is used for communication between the aforementioned terminal and other devices.
[0121] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.
[0122] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0123] In another embodiment of the present invention, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform the script scene recognition method described in any of the above embodiments.
[0124] In another embodiment of the present invention, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the script scene recognition method described in any of the above embodiments.
[0125] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0126] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0127] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0128] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention are included within the scope of protection of the present invention.
Claims
1. A method for identifying script scenes, characterized in that, include: For each scene data in the script's multiple scene data, information extraction processing is performed on the scene data based on a pre-trained recognition model to obtain scene information. The scene information includes location, or the scene information includes location and orientation, where the orientation is used to characterize the orientation information of the shooting location corresponding to the scene data at the location. If the scene information includes a location, the location is determined as the scene corresponding to the scene data; if the scene information includes both location and orientation, the scene corresponding to the scene data is determined based on the location and orientation. The recognition model includes a location recognition sub-model and a region recognition sub-model. The scene information also includes region information. The scene information is obtained by extracting information from the scene data based on the pre-trained recognition model, including: The scene data is input into the location identification sub-model for information annotation processing to obtain first annotation information, or to obtain first annotation information and second annotation information. The first annotation information is used to indicate the location in the scene data, and the second annotation information is used to indicate the orientation in the scene data. The location is extracted from the scene data based on the indication of the first annotation information, or the location and orientation are extracted from the scene data based on the indication of the first annotation information and the second annotation information. The scene data is input into the region identification sub-model for information annotation processing to obtain third annotation information, which is used to indicate the region in the scene data. Based on the indication of the third annotation information, the region information is extracted from the scene data. A target ratio corresponding to the region information is obtained, where the target ratio is the proportion of the number of second scene data to the number of first scene data. The first scene data consists of multiple scene data containing the region, and the second scene data consists of the first scene data where the region information has been successfully extracted. If the target ratio is greater than or equal to a threshold, the region information is extracted from the third scene data. If the target ratio is less than the threshold, the region information extracted from the second scene data is deleted. The third scene data consists of the first scene data where the region information has not been successfully extracted.
2. The method according to claim 1, characterized in that, The method further includes: When there is only one scene, the scene is determined as the main scene corresponding to the scene data; when there are at least two scenes, the main scene and the secondary scene corresponding to the scene data are determined based on the at least two scenes.
3. The method according to claim 1, characterized in that, The method further includes: Obtain multiple first training scenarios, each of which carries location annotation information and / or orientation annotation information; Based on the multiple first training scenarios, the location identification sub-model is iteratively trained to obtain the location identification sub-model.
4. The method according to claim 1, characterized in that, The method further includes: Multiple second training scenarios are acquired, each carrying geographic labeling information; Based on the multiple second training scenarios, the region identification sub-model to be trained is iteratively trained to obtain the region identification sub-model.
5. The method according to claim 1, characterized in that, The step of inputting the scene data into the location identification sub-model for information annotation processing to obtain first annotation information, or obtaining the first annotation information and second annotation information, includes: The geographic information in the scene data is separated to obtain the target data; The target data is input into the location identification sub-model for information annotation processing to obtain first annotation information, or the first annotation information and second annotation information.
6. The method according to claim 2, characterized in that, When the number of scenarios is one, the scenario is determined as the primary scenario corresponding to the scenario data; when the number of scenarios is at least two, after determining the primary and secondary scenarios corresponding to the scenario data based on the at least two scenarios, the method further includes: The target scene is obtained by merging and matching the main scenes corresponding to multiple scene data. If the main scene of the scene data includes the target location, the target scene in the main scene is determined as the target main scene, and the content in the main scene other than the target scene is determined as the target secondary scene.
7. The method according to claim 6, characterized in that, The step of merging and matching the main scenes corresponding to multiple scene data to obtain the target scene includes: For each main scene, determine the number of main scenes that match the main scene among the multiple main scenes corresponding to the scene data; The main scene with the most matching main scenes is identified as the target scene.
8. A device for recognizing script scenes, characterized in that, include: The first processing module is used to extract information from each scene data in the script based on a pre-trained recognition model to obtain scene information. The scene information includes location, or the scene information includes location and orientation, and the orientation is used to characterize the orientation information of the shooting position corresponding to the scene data at the location. The first determining module is configured to determine the location as the scene corresponding to the scene data when the scene information includes a location, and to determine the scene corresponding to the scene data based on the location and the orientation when the scene information includes a location and an orientation. The recognition model includes a location recognition sub-model and a region recognition sub-model, and the scene information also includes region information. The first processing module includes: The first processing unit is used to input the scene data into the location identification sub-model for information annotation processing to obtain first annotation information, or to obtain the first annotation information and second annotation information. The first annotation information is used to indicate the location in the scene data, and the second annotation information is used to indicate the orientation in the scene data. The first extraction unit is configured to extract the location from the scene data based on the indication of the first annotation information, or to extract the location and orientation from the scene data based on the indications of the first annotation information and the second annotation information; The second processing unit is used to input the scene data into the region identification sub-model for information annotation processing to obtain third annotation information, which is used to indicate the region in the scene data. The second extraction unit is used to extract the regional information from the scene data based on the indication of the third annotation information; An acquisition unit is used to acquire a target ratio corresponding to the regional information. The target ratio is the proportion of the number of second scene data to the number of first scene data. The first scene data is scene data containing the region among multiple scene data. The second scene data is the first scene data in which the regional information has been successfully extracted. The third processing unit is used to extract the regional information from the third scene data when the target ratio is greater than or equal to the threshold, and to delete the regional information extracted from the second scene data when the target ratio is less than the threshold. The third scene data is the first scene data in which the regional information was not successfully extracted.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store programs; A processor, when executing a program stored in memory, implements the steps of the method as described in any one of claims 1-7.
10. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1-7.
Citation Information
Patent Citations
Video generation method and device, video playing method and device and storage medium
CN115242980A
Method and device for establishing script scene clip video automatic extraction and retrieval by utilizing film and television works and scripts
CN116361510A