Short video production method based on artificial intelligence
By mapping and standardizing the steps of short video production in the manufacturing industry, and combining multimodal feature consistency evaluation and version verification, traceable and verifiable short videos are generated. This solves the consistency and version management problems of video production in the manufacturing industry, and achieves rapid updates and security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-03-10
AI Technical Summary
In the production of short videos for manufacturing operation guidance and equipment after-sales service, the shooting footage is affected by changes in lighting, noise, multiple people in the shot, hand obstruction, and hand tremors, making it difficult to stably edit key action segments and make it difficult to maintain a consistent sequence of steps. This can easily lead to misoperation and safety hazards. There is also a lack of verifiable version anti-mixing mechanisms and rapid update mechanisms for process changes.
By obtaining the step table and performing unified mapping and semantic normalization, structured step data is generated. Based on the structured data, the long on-site video is made usable and segmented. Multimodal feature consistency evaluation and quality penalty correction are adopted, and temporal monotonic constraints and segment uniqueness constraints are applied to solve the problem, generating target segment sequence. Version verification and difference location are performed to generate traceable short videos.
It enables the generation of standardized short videos that are traceable, verifiable, and rapidly updatable under complex field conditions, reducing the risk of misoperation, ensuring the consistency of video content with process versions, preventing the mixing of multiple versions, and supporting low-cost updates after process changes.
Smart Images

Figure CN121644930A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of short video production, and particularly relates to a short video production method based on artificial intelligence. BACKGROUND
[0002] In Chinese application No. CN202311811309.4, a short video production method based on artificial intelligence is disclosed. The method constructs a user portrait data by responding to the user's touch operation on the terminal device, and obtains historical user behavior data and user social media data corresponding to the user; obtains user real-time data through data collection of the terminal device; recommends material content data to the user portrait data and the user real-time data, and obtains material content recommendation data; responds to the user's selection operation on the material content recommendation data, and obtains material content selected data, and performs editing and synthesis on the material content selected data to obtain short video data.
[0003] Although in the field of short video production, the method can accurately understand the user's interest and preference, and realize personalized short video production, in the scene of manufacturing operation guidance and equipment after-sales short video production, due to the influence of surrounding environmental factors and human factors on the shooting materials, such as light changes, noise, multiple people entering the shot, hand occlusion and hand-held shaking, which will affect the shooting, so that it is difficult to stabilize and cut out the key action fragments, and the artificial screening is time-consuming and the consistency cannot be guaranteed, which is easy to form the situation that the picture demonstration content and the subtitle description are inconsistent, and misoperation and safety hazards are easy to form; because the video production depends on the editing experience, but the training content has long relied on word of mouth, and new people are difficult to complete standardized production; in addition, due to the formation of multiple version videos of the same operation in different equipment models and different team habits, there is a lack of unified "step-video fragment" binding mechanism, so that employees select the wrong version, and the existing technology usually regards the editing task as an isolated content editing task, lacks the mechanism of taking the "operation step table" as a strong constraint to drive fragment selection, sequence arrangement and consistency checking; at the same time, there is a lack of checkable publishing mechanism for version mixing prevention, and there is also a lack of patch type updating mechanism for process change. Therefore, a new short video production method is needed to generate traceable, checkable and quickly updated standardized short videos under complex site conditions.
[0004] In order to solve the above problems, the application provides a short video production method based on artificial intelligence. The method completes candidate fragment segmentation, matching score and constraint solving alignment by taking the operation step table as a strong constraint, and performs consistency checking, version checking, mixing prevention and patch type updating to generate traceable short videos. SUMMARY
[0005] In view of the above existing problems, the present application is proposed.
[0006] The application provides a short video production method based on artificial intelligence, aiming to solve the problems that in the manufacturing operation guidance and after-sales maintenance scene, the long video on site is affected by light fluctuation, noise, multiple people entering the camera, hand occlusion and hand-held shaking, resulting in that the key action segment is difficult to be stably intercepted and the manual screening is time-consuming; the step sequence is difficult to be strictly aligned, and misordering, missing steps and repeated segmenting are prone to occur; the inconsistency between the picture demonstration and the subtitle or dubbing is easy to cause misoperation and safety hazards; the same operation has multiple versions of multiple device models and multiple teams, and there is lack of a version mixing prevention mechanism that can be checked; and there is lack of a patch type quick updating and traceable output mechanism after process change.
[0007] To solve the above technical problems, the application provides the following technical scheme: A short video production method based on artificial intelligence, comprising: Step S1, obtaining a step table and performing unified mapping and semantic standardization processing to generate structured step data; Step S2, based on the structured step data, performing usability processing on the long video on site and executing structured segmentation to generate a candidate segment set; Step S3, based on the structured step data and the candidate segment set, performing multi-modal feature consistency evaluation and quality penalty correction on the candidate segment to calculate a matching score, and applying time monotonic constraint and segment uniqueness constraint to solve, to generate a target segment sequence; Step S4, based on the target segment sequence, performing element identification on the target segment and checking with the corresponding step elements item by item to generate a consistency result, and performing candidate segment replacement according to the consistency result to generate a shot script; Step S5, based on the shot script, performing encoding processing on the step identifier and the segment feature to generate a version check code and write it into a hidden identification mark, and performing version comparison and verification to generate a mixing prevention release piece; Step S6, based on the structured step data, performing difference positioning on the new and old steps to obtain a change step set and generating a patch piece by re-running, and outputting traceability information to generate a new version short video and traceability data.
[0008] As a preferred embodiment, the specific steps of obtaining a step table and performing unified mapping and semantic standardization processing to generate structured step data are as follows: By performing unified mapping and semantic standardization on the operation step table, the process version number, device model, step number, step description, tool list, part list and precautions are read from the step table file, the step number is executed according to the unified numbering rule, the repeated numbers and missing numbers are corrected, and the illegal characters, spaces and punctuation marks are uniformly cleaned to generate a step table standardized text.
[0009] As a preferred embodiment, the specific steps of the obtaining step table and performing uniform mapping and semantic normalization to generate structured step data further comprise: Based on step table normalized text, by establishing a standard element library and alias mapping table, synonym mapping table, and performing word segmentation and element extraction on step description, the free text is divided into four categories of fields: tool element, part element, action element, and risk element; by performing alias replacement and synonym normalization on tool names, performing uniform conversion and format unification on numerical values and units, generating element label numbers for each element and recording mapping logs before and after replacement, and generating a unique step identifier according to "process version number + step sequence number" and writing it into the record, structured step data is obtained.
[0010] As a preferred embodiment, the specific steps of the obtaining step table and performing uniform mapping and semantic normalization to generate structured step data further comprise: Based on structured step data, by performing usability processing and structured segmentation on the long video, specifically: by decoding the long video to obtain frame sequence and audio sequence, and performing audio and picture synchronization correction, performing electronic image stabilization, clarity enhancement and brightness normalization on the frame sequence, and performing noise reduction on the audio, generating usable video sequence.
[0011] As a preferred embodiment, the specific steps of the obtaining step table and performing uniform mapping and semantic normalization to generate structured step data further comprise: By calculating quality indicators according to a sliding time window, the quality indicators include jitter amplitude, blurring degree, overexposure ratio, underexposure ratio, and occlusion ratio, and when any indicator exceeds the corresponding threshold and lasts for more than the duration threshold, the time period is determined as an invalid segment and is removed, while the occlusion penalty value and the jitter penalty value of the remaining segment are calculated; after removing the invalid segment, based on the action change boundary, tool entering and leaving boundary, and voice instruction boundary, a set of cut points is generated, and the lower limit duration constraint and adjacent cut point merging rule are applied to complete the cutting, the key frame of each candidate segment is extracted, the key frame fingerprint is generated, the voice text and recognition element index are saved, and the candidate segment set is generated.
[0012] As a preferred embodiment, the specific steps of the obtaining step table and performing uniform mapping and semantic normalization to generate structured step data further comprise: Based on the candidate segment set, a matching score value is calculated by using a multi-modal consistency evaluation and quality penalty correction method on the structured step data and the candidate segment set, and a target segment sequence is generated by constraint solving, wherein the multi-modal consistency evaluation includes: obtaining an action consistency score according to comparison of an action recognition result and a step action element, obtaining a tool consistency score according to comparison of a tool recognition result obtained by target detection and a step tool element and a part element set, obtaining a voice prompt consistency score according to whether a step number word or an instruction word is contained in a voice recognition text, and obtaining a basic consistency score by normalizing the three types of scores and fusing them according to a pre-set weight.
[0013] As a preferred embodiment, the specific steps of the method for calculating the matching score of the candidate segment by using the multi-modal feature consistency evaluation and quality penalty correction method and imposing time monotonicity constraint and segment uniqueness constraint to generate the target segment sequence further include: After obtaining the basic consistency score, in order to enable the quality degradation factors of field occlusion and handheld jitter to participate in segment screening in a calculable manner, a quality penalty term is further calculated for each candidate segment, and the quality penalty term and the basic consistency score are combined to generate a matching score value, forming a matching score value matrix for constraint solving, wherein the quality penalty term at least includes an occlusion penalty value and a jitter penalty value, and the algorithm formula for calculating the matching score value is: , wherein Point is the matching score value of the step and the candidate segment, i is the index number of the matching step, j is the index number of the candidate segment, is the weight coefficient of the action consistency score, is the weight coefficient of the tool consistency score, is the weight coefficient of the voice prompt consistency score, is the occlusion penalty coefficient, is the jitter penalty coefficient, is the action consistency score, is the tool consistency score, is the voice prompt consistency score, is the occlusion index of the candidate segment, is the jitter index of the candidate segment; After the matching score value is calculated, a lower limit score threshold minPoint is set, when Point(i,j) < minPoint, step i and candidate segment j are determined as unassignable pair, and the unassignable pair is removed from the optional set of constraint solving to form an unassignable gating mechanism; when Point(i,j) ≥ minPoint, step i and candidate segment j are determined as assignable pair, and the assignable pair is included in the optional set of constraint solving; the quality penalty correction includes: the general occlusion penalty value and the jitter penalty value deduct points from the low-quality segment to form the matching score value matrix, while retaining the candidate replacement list sorted by score, and generating the target segment sequence.
[0014] As a preferred embodiment, the specific steps of taking element identification on the target segment and checking with the corresponding step element one by one to generate consistency results, and performing candidate segment replacement according to the consistency results to generate the shot script are: Based on the target segment sequence, element identification is performed on each target segment to respectively output tool identification results, part identification results, action identification results and risk action identification results, and the identification results are checked with the tool elements, part elements, action elements and risk elements of the step one by one to generate consistency results; the consistency results include consistency determination value and inconsistent field list, when the field determination value is higher than the consistency threshold, it is determined that the field is consistent and the field is removed from the inconsistent field list; when the consistency determination value of any field is lower than the consistency threshold, the target segment is marked as inconsistent segment, and the candidate segments are taken from the candidate replacement list in order according to the matching score value from the step to recalculate the consistency results, and the candidate segment is selected for replacement, while updating the step-segment mapping relationship and the corresponding time code to generate the shot script.
[0015] As a preferred embodiment, the specific steps of taking encoding processing on step identification and segment characteristics to generate version check code and write into hidden identification mark, and performing version comparison verification to generate anti-mixing release finished piece are: The tool detection result in the target segment is counted to obtain a tool category value, a start and end timestamp of the target segment is calculated to obtain a segment duration value, the segment fingerprint value, the tool category value and the segment duration value are spliced in a preset field order and subjected to a hash operation to generate a segment feature code corresponding to the operation step; a version check code is generated by performing an encoding operation based on the process version number, the step unique identification sequence and the segment feature code sequence, and the version check code is written into the hidden identification mark; the frequency domain transformation is performed on the published key frame, and the version check code is embedded in the preset intermediate frequency coefficient position; the version check code is extracted from the hidden identification mark in the distribution link, and the current effective process version number and the corresponding shot script are used to recalculate the current check code; when the extracted version check code is consistent with the current check code, an allowed playing result is output, and when the two are inconsistent, a playing blocking or distribution blocking result is output and the inconsistency reason is recorded, and the anti-mixed published segment and the version check record are generated.
[0016] As a preferred embodiment, the specific steps of taking difference positioning on the new and old steps to obtain a change step set and generating a patch segment by re-running, while outputting trace information, to generate a new version short video and trace data are as follows: By monitoring the process version number in real time, when the process version number is detected to change, difference positioning is performed on the new and old structured step data, and the step number, the action element set, the tool element set, the part element set and the risk element set are compared item by item to generate a change step set and output a change reason field; the unchanged steps are directly reused to the existing target segment and shot script record, and the change steps are re-run according to the processing link of “candidate segment retrieval-matching score calculation-constraint solving-consistency checking and replacement-version check code updating”, wherein the candidate segment retrieval is preferentially selected from the existing candidate segment set according to the tool element and the action element, and is sorted according to the matching score value; when the filtered candidate segment is empty, a to-be-supplemented recording prompt is output; then the reused segment and the re-run segment are synthesized into a patch segment according to the step order, and the hidden identification mark and the version check code are updated synchronously to obtain a new version short video, and the trace information is output. Beneficial effects
[0017] 1. By taking multi-modal consistency scoring on the candidate segment and superimposing occlusion jitter penalty, and then applying time monotonicity and segment uniqueness constraint solving, a target segment sequence stably aligned according to the step order is generated to obtain an “step table driven interpretable alignment result”, thereby reducing the risk of misoperation caused by disorder, mismatch and repeated selection of segments.
[0018] 2. By taking "step identification + segment features" coding generation version check code and hidden write, and executing version comparison output at the playback end, the generation of a checkable anti-mixing distribution mechanism is allowed or prevented, thereby preventing the spread and playback of old and wrong versions of videos, and ensuring that the on-site use content is consistent with the version of the process.
[0019] 3. By taking the difference positioning of new and old steps to obtain a set of changed steps, and only performing candidate segment generation, matching alignment, consistency error correction and check writing on the changed steps, patching and tracing data are generated, thereby realizing low-cost and rapid updating in the process of frequent process changes, and providing a traceable link of steps-segments-versions for auditing and review. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 Flowchart of the present application.
[0021] Figure 2 Technical effect comparison chart of the present application, in which the black column chart is the present application and the gray column chart is the prior art. DETAILED DESCRIPTION
[0022] In order to make the technical means, creative features, purposes and effects achieved by the present application easy to understand, the following specific embodiments are further described, but the following embodiments are only preferred embodiments of the present application, not all. Based on the embodiments in the embodiments, other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application. Unless otherwise specified, the experimental methods in the following embodiments are conventional methods, and the materials, reagents, etc. used in the following embodiments can be obtained from commercial channels unless otherwise specified.
[0023] Example 1, in combination Figure 1 A flowchart of a short video production method based on artificial intelligence is shown as follows, which is the specific implementation steps: Step S1, obtain the step table and perform unified mapping and semantic standardization processing to generate structured step data; In step S1, the job step data is converted into structured step data that can be calculated and aligned, so that the step matching deviation caused by different personnel expression differences can be solved, and the unified standard can be used as a reference for subsequent scoring and checking. Specifically, by unified mapping and semantic standardization of the operation step table, the process version number, equipment model, step number, step description, tool list, part list and matters needing attention are first read from the step table file, and unified numbering rules are performed on the step number to correct repeated numbers and missing numbers, and unified cleaning is performed on illegal characters, spaces and punctuation to generate step table standardized text; based on the step table standardized text, by establishing a standard element library and an alias mapping table and a synonym mapping table, and performing word segmentation and element extraction on the step description, the free text is divided into four types of fields: tool elements, part elements, action elements and risk elements; by performing alias replacement and synonym normalization on the tool name, performing unified conversion and format unification on the numerical value and unit, generating element label number for each element and recording mapping log before and after replacement, and generating step unique identification according to "process version number + step number" and writing into the record, structured step data is obtained, thereby providing strong constraint input for subsequent segmentation, matching and consistency checking; The step table standardized text is formed by splicing the step number, step description, tool list, part list and matters needing attention of each step according to the preset field order, and is used as the input for subsequent alias mapping, synonym normalization, element extraction and structured step data generation.
[0024] In step S2, based on the structured step data, the long video on site is processed and structured segmentation is performed to generate a candidate segment set; In step S2, the long video on site is converted into a set of available and searchable candidate segments, which is used to eliminate invalid content and form a segment unit that can be selected by an algorithm, thereby reducing noise and complexity, and reducing the search space for subsequent alignment and error correction; Specifically, based on structured step data, usability processing and structured segmentation are performed on long-form video footage. Specifically, the long video is decoded to obtain frame and audio sequences, and audio-visual synchronization correction is performed. Electronic image stabilization, sharpness enhancement, and brightness normalization are applied to the frame sequences, while noise reduction is performed on the audio to reduce the impact of handheld shaking, lighting fluctuations, and ambient noise. This generates a stable and clear usable video sequence, enabling subsequent action recognition and reducing segmentation misjudgments and matching score deviations caused by image shaking and noise. Quality indicators are calculated using a sliding time window, including shaking amplitude, blur, overexposure ratio, and underexposure. The system calculates the exposure ratio and occlusion ratio, and when any indicator exceeds the corresponding threshold and continues to exceed the duration threshold, the time period is judged as an invalid segment and removed. At the same time, the occlusion penalty value and jitter penalty value of the retained segment are calculated. After removing invalid segments, a set of segmentation points is generated based on the action change boundary, tool entry and exit boundary and voice command boundary. The lower limit duration constraint and adjacent segmentation point merging rule are applied to the segmentation points to complete the segmentation. The key frame of each candidate segment is extracted, the key frame fingerprint is generated, and the voice text and recognition element index are saved to generate a set of candidate segments. This enables the transformation of complex on-site materials into fragmented inputs that can be matched and quantified for evaluation. The available video sequence is a set of frame sequences corresponding to video time segments that meet preset quality conditions, obtained by performing usability processing on the long video, along with their corresponding audio sequences. The usability processing includes image stabilization, sharpness enhancement, brightness normalization, and audio noise reduction, and invalid time segments that do not meet the quality conditions are removed based on jitter amplitude thresholds, occlusion ratio thresholds, and sharpness thresholds. The available video sequence is formed by splicing the retained time segments in their original chronological order and provides input for subsequent structured segmentation and candidate segment generation. The long video includes a video sequence and an audio sequence; wherein the video sequence consists of a sequence of frames arranged in chronological order, and each frame is assigned a corresponding timestamp; the audio sequence consists of audio samples arranged in chronological order, and is aligned with the frame sequence through timestamps; the long video also includes shooting parameter information, which includes at least frame rate, resolution, and encoding format; The departure boundary is determined by performing frame-by-frame detection on a preset object, such as a tool or part. When the detection state of the object changes from present to non-present and remains so for no less than a preset number of frames, the timestamp corresponding to the frame in which the state first changes to non-present is determined as the departure boundary. The presence of the detection state is determined by a detection confidence level that is not lower than a preset confidence threshold. The voice instruction boundary is obtained by voice recognition on the audio sequence to obtain a time-stamped word sequence, and when a keyword in the word sequence completely matches a step number word, an instruction word, or a tool / part name word, the starting time stamp of the keyword is determined as the voice instruction boundary; wherein the recognition confidence of the keyword is not less than a preset voice confidence threshold; The lower limit time length constraint and the adjacent split point merging rule arrange the split point sequence in time sequence as {a1, a2, …, an}, when the interval (ak+1-ak) between adjacent split points is less than L_min, the last split point ak+1 is deleted, and [ak, ak+2] is determined as the time range of the merged candidate segment; sequentially iterate until all adjacent split point intervals are not less than L_min, thereby obtaining the candidate segment set satisfying the lower limit time length constraint, wherein a is the split point time stamp, the split point time stamp represents the segmentation boundary position on the long video time axis, k is the split point sequence number and k is a positive integer and satisfies 1≤k≤n−1, n is the total number of split points, and n is a positive integer, L_min is the lower limit time length threshold, used to limit the minimum length of the candidate segment; Step S3, based on the structured step data and the candidate segment set, a method of multi-modal feature consistency evaluation and quality penalty correction is used to calculate the matching score, and time monotonicity constraint and segment uniqueness constraint are applied to solve, to generate the target segment sequence; Wherein, step S3 establishes a stable one-to-one correspondence between the step and the candidate segment, which plays a core alignment role, thereby avoiding out-of-order, wrong selection, and repeated selection, and suppressing occlusion and jitter segments from entering the film; Specifically, based on the candidate segment set, a multi-modal consistency evaluation and quality penalty correction method is used to calculate the matching score value of the structured step data and the candidate segment set, and the target segment sequence is generated by constraint solving, wherein the multi-modal consistency evaluation includes: obtaining an action consistency score by comparing the action recognition result with the step action element, obtaining a tool consistency score by comparing the tool recognition result obtained by target detection with the step tool element and part element set, obtaining a voice prompt consistency score by whether the step number word or the instruction word is contained in the voice recognition text, and normalizing the three types of scores and fusing them according to the pre-set weight to obtain the basic consistency score; after obtaining the basic consistency score, to enable quality degradation factors such as on-site occlusion and handheld jitter to participate in segment selection in a calculable manner, further calculate a quality penalty term for each candidate segment, and combine the quality penalty term with the basic consistency score to generate a matching score value, thereby forming a matching score value matrix for constraint solving, wherein the quality penalty term at least includes an occlusion penalty value and a jitter penalty value, and the algorithm formula for calculating the matching score value is: , wherein Point is the matching score value of the step and the candidate segment, i is the index number of the matching step, j is the index number of the candidate segment, is a weight coefficient of the action consistency score, is a weight coefficient of the tool consistency score, is a weight coefficient of the voice prompt consistency score, is a blocking penalty coefficient, is a jitter penalty coefficient, is the action consistency score, is the tool consistency score, is the voice prompt consistency score, is the blocking indicator of the candidate segment, is the jitter indicator of the candidate segment; After the matching score value is calculated, a lower limit score threshold minPoint is set, when Point(i,j) < minPoint, the step i and the candidate segment j are determined as an unassignable pair, and the unassignable pair is removed from the optional set of constraint solving to form an unassignable gating mechanism; when Point(i,j) ≥ minPoint, the step i and the candidate segment j are determined as an assignable pair, and the assignable pair is included in the optional set of constraint solving, which is used for subsequent assignment solving under the time monotonic constraint and the segment unique constraint, to form a gating screening mechanism of the matching score value matrix; In the manufacturing operation guidance and after-sales maintenance training scene, the actual application process of the algorithm is that the system first obtains the action points, tool elements and step orders of each step from the step table, and simultaneously cuts the long video on site to obtain a candidate segment set sorted by time; the system respectively identifies the action label, tool label and voice order text of each candidate segment, and calculates the action consistency score, tool consistency score and order hit score with each step; the system simultaneously calculates the blocking indicator and jitter indicator of the candidate segment, and adds the blocking and jitter as penalty items to the matching score; when the matching score is lower than the minimum threshold, the "step-segment" combination is determined as unassignable, which is directly prohibited from entering the subsequent solving; in the assignable combination, the system applies the time monotonic constraint (the step order is consistent with the segment time order) and the segment unique constraint (the same segment is not repeated assignment), to solve the maximum total matching score as the target, and output the target segment sequence corresponding to each step, so that when there is light fluctuation, noise, multiple people in the shot, hand blocking and handheld jitter on site, the algorithm can still automatically align the long video to the short segment sequence with correct step order, and reduce the low-quality segment and error segment into the film, thereby reducing the risk of manual finding and mismatching; The quality penalty correction includes: a general occlusion penalty value and a jitter penalty value deduct points from low-quality segments to form a matching score value matrix, and segments below a minimum score threshold are set as unassignable, a time monotonic constraint and a segment uniqueness constraint are imposed on the assignment of "step-segment", and a dynamic programming recursive solution is performed with the maximum total matching score value as the target while constraining the minimum segment length; when there is no assignable segment for a certain step, mark it as empty and output a prompt for missing recording, while retaining a candidate replacement list sorted by score, generating a target segment sequence, so as to avoid misordering, missing steps and repeated segmenting and improve the stability of the aligned segment sequence.
[0025] Step S4, based on the target segment sequence, performing element recognition on the target segment and checking with the corresponding step element by item, generating a consistency result, and performing candidate segment replacement according to the consistency result to generate a shot script; Step S4 is used to ensure that the picture demonstration is consistent with the step explanation, and plays a role in quality control and error correction closed loop, and outputs a consistency result through identification and checking, and triggers a replacement marker, causing the risk of misoperation caused by picture A and explanation B, where A is the action of using a common wrench to directly tighten the connecting bolt in the picture, and B is the explanation text (or voiceover); Specifically, based on the target segment sequence, element recognition is performed on each target segment to output tool recognition results, part recognition results, action recognition results, and risk action recognition results, and the recognition results are checked with the tool elements, part elements, action elements, and risk elements of the step one by one, thereby generating a consistency result, the consistency result includes a consistency determination value and an inconsistency field list, when the field determination value is higher than the consistency threshold, it is determined that the field is consistent and the field is excluded from the inconsistency field list; when the consistency determination value of any field is lower than the consistency threshold, the target segment is marked as an inconsistent segment, and candidate segments are taken from the candidate replacement list output by step S3 in the order of matching score value, and the consistency result is recalculated, the candidate segment that meets the consistency threshold and has the largest matching score value is selected to replace the current inconsistent segment, and the step-segment mapping relationship and the corresponding time code are updated; when there is no candidate segment that meets the consistency threshold, output a missing recording prompt and retain a manual review marker to generate a shot script, the shot script includes a step unique identifier, a segment start and end time, a subtitle text, a voiceover text, a part prompt, and a caution prompt, thereby ensuring that the film content is consistent with the step table item by item and reducing the cost of manual proofreading; The element recognition is performed by respectively performing recognition processing on the picture frame sequence and the audio sequence of the target segment to output element results for consistency checking; wherein target detection and action recognition are performed on the picture frame sequence to obtain tool element recognition results, part element recognition results and action element recognition results, and a corresponding timestamp and confidence are given for each recognized element; speech recognition is performed on the audio sequence to obtain step order text and its appearance timestamp; the above recognition results are written into an element recognition result set according to a preset field, for item-by-item checking with tool elements, part elements and action elements in the structured step data.
[0026] In step S5, based on the shot list script, step identification and segment features are coded, version check codes are generated and written into hidden identification marks, and version comparison and checking are performed to generate anti-mixing publishing films; In step S5, the film is bound with the process version and the publishable checking ability is provided, which can prevent mixing use and prevent the spread of old versions. By hiding the check code in the playback terminal and comparing, the blocking and prompting of wrong and old versions are realized, and the correct version of the video used in the field is ensured. Specifically, based on the shot list script, a segment feature coding operation is performed on each job step corresponding to the target segment in the shot list script, the segment feature coding operation includes: a key frame set is extracted at a preset time interval within the time range of the target segment, the frame fingerprint value of each key frame in the key frame set is calculated, the frame fingerprint values are aggregated to obtain a segment fingerprint value, the tool category with the highest appearance frequency in the tool detection results within the target segment is counted and determined as the tool category value, the start and end timestamps of the target segment are calculated to obtain a segment duration value, and the segment fingerprint value, the tool category value and the segment duration value are spliced in a preset field order and subjected to a hash operation, thereby generating a segment feature code corresponding to the job step. Based on the process version number, the step unique identification sequence and the segment feature code sequence, a version check code is generated by performing coding operation, and the version check code is written into the hidden identification mark. The key frames of the published film are subjected to frequency domain transformation, and the version check code is embedded in the preset intermediate frequency coefficient position, so that the hidden identification mark can still be extracted after transcoding. The hidden identification mark is extracted to obtain the version check code in the distribution link, and the current effective process version number and the corresponding shot list script are used to recalculate the current check code. When the extracted version check code is consistent with the current check code, the playback permission result is output, and when they are inconsistent, the playback blocking or distribution blocking result is output and the inconsistency reason is recorded. The anti-mixing publishing film and the version check record are generated, so that the film version can be checked, blocked and traced, thereby avoiding training deviation and safety risks caused by multiple versions.
[0027] Step S6, based on the structured step data, the difference positioning is taken on the new and old steps to obtain a change step set and generate a patch, and meanwhile, traceability information is outputted, and a new version of short video and traceability data are generated; Among them, step S6 is used for quickly updating and forming a traceable link after process change; the low-cost iteration and auditability are realized, the change step is generated by difference positioning and re-calculation, the patch is generated, and the traceability information is outputted, so as to reduce the full re-production cost and support subsequent review and accountability; Specifically, by monitoring the process version number in real time, when the process version number is detected to change, difference positioning is performed on new and old structured step data, and step number, action element set, tool element set, part element set and risk element set are compared item by item respectively, a change step set is generated and a change reason field is outputted; by directly reusing the existing target segment and shot script record for unchanged steps, the change steps are re-run according to the processing link of “candidate segment retrieval-matching score calculation-constraint solving-consistency checking and replacement-version check code updating”, wherein the candidate segment retrieval is preferentially selected from the existing candidate segment set according to the tool element and the action element, and is sorted according to the matching score value, and when the filtered candidate segment is empty, a to-be-supplemented recording prompt is outputted; then, the reused segment and the re-run segment are synthesized into a patch according to the step order, and the hidden identification mark and the version check code are updated synchronously, a new version of short video is obtained, traceability information is outputted, and the traceability information includes new and old process version numbers, change step set, step-segment mapping relationship, replacement log, version check code and generation time, so that the traceability and checkability are ensured, the rapid iteration and update are realized, and the training deviation risk caused by version switching is reduced; In combination Figure 2 As shown in the technical effect comparison chart of the short video production method based on artificial intelligence, the black column chart is the technical effect of the present application, and the gray column chart is the prior art, wherein Figure 2 It can be seen that the technical effect of the present application is better than that of the prior art.
[0028] Embodiment 2, a short video production method based on artificial intelligence based on embodiment 1, the specific scheme is as follows: In this embodiment, the application scenario is “short video production for equipment after-sales maintenance training”: the maintenance personnel receive the work order, the work order contains the equipment model, the fault code and the version number, and the on-site long video and the electric tool event log (containing the time stamp and the torque reaching the standard state) are collected during the maintenance process, and the system automatically generates a short video which is consistent with the maintenance process, can be checked to prevent misuse, and can be patched and updated.
[0029] The step S1 generates structured step data by performing association extraction, field unification and semantic standardization processing on the work order and the standard operation step table; the association extraction includes: locating the corresponding maintenance process in the step library according to the work order fault code, and determining the maintenance step sequence; the field unification includes: mapping the equipment model, version number, tool name and part name to standard codes according to a preset dictionary, and unifying the unit and format to a preset format; the semantic standardization includes: by decomposing the action description of each step according to “object + method + result”, generating action keywords and writing them into the step field, and generating a step unique identifier according to “process version number + step sequence number” and writing it into the record, to obtain structured step data, thereby providing strong constraint input that can be directly calculated for subsequent segmentation, alignment and consistency checking.
[0030] The step S2 generates a candidate segment set by performing usability processing and structured segmentation on the long video on site based on the step data; the usability processing includes: performing image stabilization, clarity enhancement and brightness unification on the video, and eliminating invalid segments that meet any of the following conditions: no action change segment with motion energy less than a motion threshold, occlusion segment with an occlusion ratio not less than an occlusion threshold and a duration not less than a duration threshold, and low clarity segment with a clarity index less than a clarity threshold; the structured segmentation includes: performing tool detection on the picture and generating a segmentation boundary when the tool category appears, disappears or switches, performing speech recognition on the audio and generating a segmentation boundary when a step password appears, and aligning the electric tool event log to the video time axis according to the timestamp, and generating a segmentation boundary when an event such as “torque meets the standard / fastening is completed / dismantling is completed” appears; then, the various boundaries are merged, and the minimum segment duration threshold and the maximum segment duration threshold are applied to the adjacent boundaries, to obtain the candidate segment set, thereby being able to convert complex field materials into retrievable and replaceable segment units.
[0031] Step S3 is based on the step data and the candidate segments, and a matching score is calculated by taking a multi-modal feature consistency evaluation and a quality penalty correction method on the candidate segments, and a time monotonic constraint and a segment unique constraint are solved to generate a target segment sequence; the multi-modal features include: action labels obtained by action recognition, tool labels obtained by target detection, step order text obtained by speech recognition, and "torque value / standard reaching state" labels obtained by tool event logs; the consistency evaluation includes: calculating action consistency scores, tool / part consistency scores, order hit scores, and torque event consistency scores, and obtaining a basic consistency score by fusing them according to a preset weight; the quality penalty correction includes: deducting the matching score based on the occlusion index and the jitter index, and setting the candidate segments below the minimum score threshold as unassignable; under the time monotonic constraint and the segment unique constraint, sequence alignment solving is performed to obtain a one-to-one corresponding target segment sequence in step order, so that stable alignment of maintenance steps and video segments can be realized under the conditions of occlusion, jitter and multiple people entering the shot, and error segments and low-quality segments are prevented from entering the film.
[0032] Step S4 is based on the target segment sequence, and a consistency result is generated by taking an element recognition and item-by-item checking method on the target segments, and a segment replacement processing is performed according to the consistency result to generate a shot script; the item-by-item checking at least includes: tool consistency checking, part consistency checking, action consistency checking and torque event consistency checking; the consistency result includes a consistency determination value and an inconsistency field list: when all field determination values are not lower than the consistency threshold, the consistency determination value is set to consistent and the target segment is kept unchanged; when there is any field determination value lower than the consistency threshold, the consistency determination value is set to inconsistent and the inconsistency field list is output; in the inconsistent case, the target segment is replaced in order from high to low according to the candidate segment matching score of this step and rechecked; when the replacement times reach the preset upper limit and still inconsistent, a review mark is generated and written into the shot script, and the shot script is generated, so that the risk of misoperation caused by "picture demonstration and order / instruction inconsistency" can be avoided.
[0033] Step S5 generates a version check code and writes it into a hidden identification mark based on the split-screen script by taking encoding processing on the step identification, segment feature and fault code, and generates a release piece for anti-mixing use by performing version comparison at the playing end; the segment feature code contains tool / part label summary, key frame hash and segment duration mark; the version check code binds the work order fault code, device model, process version number, step unique identification sequence and segment feature sequence, and generates a unique check result through a preset encoding rule; the playing end reads the hidden identification mark to extract the version check code, and compares it with the corresponding check code in the version library according to the "fault code-device model-process version", and outputs allowed playing if the comparison is consistent, and outputs prevented playing and prompts version mismatch if the comparison is inconsistent, so as to prevent the use of old version process, wrong version process and mismatched fault code video, generate a release piece for anti-mixing use, and thus the bidirectional checkable of the publishing side and the using side can be realized.
[0034] Step S6 generates a patch piece and outputs trace information based on the step data by taking difference positioning on the new and old steps to obtain a change step set, and re-executes steps S2 to S5 of embodiment 2 only for the change step set to generate a patch piece, and generates a new version short video and trace data; the trace information at least includes: device model, step-segment start and end time, torque standard result, consistency state, version check code and generation time, and is used for audit review and responsibility trace, so that only the affected steps are recalculated after the process or flow is changed, and then the re-shooting and re-cutting are reduced and the update cycle is shortened.
[0035] Through the above embodiment 2, the step alignment, error correction closed loop, version anti-mixing use and change patch update are realized in the "work order driven maintenance process" scene, and additionally, the tool event log is included in the segmentation and consistency check, so that the piece meets the "correct picture, correct step, data standard and correct version" at the same time.
[0036] The basic principles, main features and advantages of the present application are shown and described above. It should be understood by those skilled in the art that the present application is not limited by the above embodiments, and the above embodiments and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. An artificial intelligence-based short video production method, characterized in that: Step S1, obtain the step table and perform unified mapping and semantic standardization processing to generate structured step data; Step S2, based on the structured step data, perform usability processing on the long video on site and execute structured segmentation to generate a candidate segment set; Step S3, based on the structured step data and the candidate segment set, perform multi-modal feature consistency evaluation and quality penalty correction on the candidate segments to calculate the matching score, and apply time monotonicity constraint and segment uniqueness constraint to solve, to generate a target segment sequence; Step S4, based on the target segment sequence, perform element recognition on the target segment and check it with the corresponding step element item by item to generate a consistency result, and perform candidate segment replacement according to the consistency result to generate a split script; Step S5, based on the split script, perform encoding processing on the step identifier and segment feature to generate a version check code and write it into a hidden identification mark, and perform version comparison and verification to generate a mixed-use prevention release finished piece; Step S6, based on the structured step data, perform difference positioning on the new and old steps to obtain a change step set and generate a patch finished piece by re-running, while outputting traceability information, to generate a new version short video and traceability data. 2.The short video production method based on artificial intelligence of claim 1, wherein: The specific steps of obtaining the step table and performing unified mapping and semantic standardization processing to generate structured step data are: By performing unified mapping and semantic standardization on the job step table, first read the process version number, equipment model, step number, step description, tool list, part list and precautions from the step table file, and perform unified numbering rules on the step number to correct repeated numbers and missing numbers, and simultaneously perform unified cleaning on illegal characters, spaces and punctuation to generate step table standardized text. 3.The short video production method based on artificial intelligence of claim 2, wherein: The specific steps of obtaining the step table and performing unified mapping and semantic standardization processing to generate structured step data further include: Based on the step table standardized text, by establishing a standard element library as well as an alias mapping table and a synonym mapping table, and performing word segmentation and element extraction on the step description, the free text is split into four types of fields: tool elements, part elements, action elements and risk elements; by performing alias replacement and synonym normalization on tool names, performing unified conversion and format unification on numerical values and units, simultaneously generating element tag numbers for each element and recording mapping logs before and after replacement, and generating step unique identifiers according to process version number + step number and writing them into records, structured step data is obtained.
4. The short video production method based on artificial intelligence according to claim 1, characterized in that: The specific steps of performing usability processing on the long video on site and executing structured segmentation to generate a candidate segment set are: Based on the structured step data, perform usability processing and structured segmentation on the long video on site, specifically: decode the long video to obtain frame sequence and audio sequence, perform audio and picture synchronization correction, perform electronic image stabilization, clarity enhancement and brightness normalization on the frame sequence, and perform noise reduction on the audio to generate usable video sequence.
5. The short video production method based on artificial intelligence according to claim 4, characterized in that: The specific steps of performing usability processing on the long video on site and executing structured segmentation to generate a candidate segment set are: The quality indicators include jitter amplitude, blurring degree, overexposure proportion, underexposure proportion and occlusion proportion, and when any indicator exceeds the corresponding threshold and lasts for more than a duration threshold, the time period is determined as an invalid segment and is removed, while the occlusion penalty value and the jitter penalty value of the remaining segment are calculated; after removing the invalid segment, a set of segmentation points is generated based on the action change boundary, the tool entering and leaving boundary and the voice instruction boundary, and the segmentation points are subjected to a lower limit duration constraint and an adjacent segmentation point merging rule to complete segmentation, extract the key frame of each candidate segment, generate a key frame fingerprint, save the voice text and the recognition element index, and generate a candidate segment set.
6. The short video production method based on artificial intelligence according to claim 1, characterized in that: The specific steps of the method for taking multi-modal feature consistency evaluation and quality penalty correction on the candidate segments to calculate a matching score and solve the constraints to generate a target segment sequence are as follows: Based on the candidate segment set, a multi-modal consistency evaluation and quality penalty correction method is used to calculate a matching score value, and a target segment sequence is generated by constraint solving, wherein the multi-modal consistency evaluation includes: obtaining an action consistency score according to the comparison of the action recognition result and the step action element, obtaining a tool consistency score according to the comparison of the tool recognition result obtained by target detection and the step tool element and part element set, and obtaining a voice prompt consistency score according to whether the step number word or the instruction word is contained in the voice recognition text, and then the three types of scores are normalized and fused according to the pre-set weight to obtain a basic consistency score.
7. The short video production method based on artificial intelligence according to claim 6, characterized in that: The specific steps of the method for taking multi-modal feature consistency evaluation and quality penalty correction on the candidate segments to calculate a matching score and solve the constraints to generate a target segment sequence are as follows: After obtaining the basic consistency score, in order to make the quality degradation factors of field occlusion and handheld jitter participate in the segment screening in a calculable way, a quality penalty term is further calculated for each candidate segment, and the quality penalty term and the basic consistency score are combined to generate a matching score value, forming a matching score value matrix for constraint solving, and the quality penalty term at least includes an occlusion penalty value and a jitter penalty value, wherein the algorithm formula for calculating the matching score value is: , wherein Point is a matching score value of a step with a candidate segment, i is an index number of a matching step, j is an index number of a candidate segment, is a weight coefficient of the action consistency score, is a weight coefficient of the tool consistency score, is a weight coefficient of the voice prompt consistency score, is an occlusion penalty coefficient, is a jitter penalty coefficient, is the action consistency score, is the tool consistency score, is the voice prompt consistency score, is an occlusion indicator of the candidate segment, is a jitter indicator of the candidate segment; After calculating the matching score value, a lower limit score threshold minPoint is set, when Point(i,j)<minPoint, the step i and the candidate segment j are determined as an unassignable pair, and the unassignable pair is removed from the selectable set for constraint solving, to form an unassignable gating mechanism; when Point(i,j)≥minPoint, the step i and the candidate segment j are determined as an assignable pair, and the assignable pair is included in the selectable set for constraint solving; the quality penalty correction includes: the general occlusion penalty value and the jitter penalty value deduct points from the low-quality segment to form the matching score value matrix, while retaining a candidate replacement list sorted by score, and generating a target segment sequence. 8.The short video production method based on artificial intelligence of claim 1, wherein: The specific steps for taking element identification on the target segment and checking with corresponding step elements item by item to generate consistency results and performing candidate segment replacement according to the consistency results to generate the shot script are: Based on the target segment sequence, element identification is performed on each target segment to respectively output tool identification results, part identification results, action identification results and risk action identification results, and the identification results are checked with tool elements, part elements, action elements and risk elements of the step item by item to generate consistency results; the consistency results include consistency judgment values and inconsistent field lists, when the field judgment value is higher than the consistency threshold value, it is judged that the field is consistent and the field is excluded from the inconsistent field list; When the consistency judgment value of any field is lower than the consistency threshold value, the target segment is marked as an inconsistent segment, and the candidate segments are taken from the candidate replacement list in order according to the matching score value from the step to recalculate the consistency results and select the candidate segments for replacement, while updating the step-segment mapping relationship and the corresponding time code to generate the shot script. 9.The short video production method based on artificial intelligence of claim 1, wherein: The specific steps for taking encoding processing on the step identification and segment features to generate version check codes and write them into hidden identification marks, and performing version comparison and verification to generate anti-mixing release finished pieces are: By extracting a key frame set in the time range of the target segment at a preset time interval, calculating the frame fingerprint value of each key frame in the key frame set, and aggregating the frame fingerprint values to obtain a segment fingerprint value, the occurrence times of tool categories in the tool detection results in the target segment are counted and determined as tool category values, the start and end timestamps of the target segment are calculated to obtain a segment duration value, and the segment fingerprint value, tool category value and segment duration value are spliced in a preset field order and subjected to a hash operation to generate a segment feature code corresponding to the work step; based on the process version number, step unique identification sequence and segment feature code sequence, a version check code is generated by performing encoding operation, and the version check code is written into the hidden identification mark; the key frames of the released finished pieces are subjected to frequency domain transformation, and the version check code is embedded in the preset intermediate frequency coefficient position; in the distribution link, the hidden identification mark is extracted to obtain the version check code, and based on the currently effective process version number and the corresponding shot script, the current check code is recalculated; when the extracted version check code is consistent with the current check code, an allowed playing result is output, and when they are inconsistent, a blocked playing or blocked distribution result is output and the inconsistent reason is recorded to generate the anti-mixing release finished pieces and the version check record. 10.The short video production method based on artificial intelligence of claim 1, wherein: The specific steps for taking difference positioning on the new and old steps to obtain a change step set and performing re-running to generate a patch finished piece, while outputting trace information, to generate a new version short video and trace data are: By real-time monitoring of the process version number, when the process version number change is detected, the difference positioning is performed on the new and old structured step data, the step number, action element set, tool element set, part element set and risk element set are compared item by item respectively, the change step set is generated and the change reason field is output; by directly reusing the existing target fragment and the shot script record of the unchanged step, the change step is re-run according to the processing link of "candidate fragment retrieval-matching score calculation-constraint solving-consistency check and replacement-version check code updating", wherein the candidate fragment retrieval is preferentially selected in the existing candidate fragment set according to the tool element and the action element, and is sorted according to the matching score value, and when the filtered candidate fragment is empty, the to-be-supplemented recording prompt is output; then the reused fragment and the re-run fragment are synthesized into a patch according to the step order, and the hidden identification mark and the version check code are updated synchronously, the new version short video is obtained, and the traceability information is output.
Citation Information
Patent Citations
Short video production method and system based on artificial intelligence
CN117714797A