Image recognition processing method and system applied to video conversion
By presetting video slicing standards and benchmark image judgment, the problem of error recognition in video conversion is solved, and quick and accurate error marking and repair are achieved.
Patent Information
- Application Number
- CN202510295461.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to automatically identify the wrong position during video conversion, resulting in problems such as garbled code or misalignment of video clips.
By presetting the video slicing standards, the reference image is intercepted and the difference between the front and back images is judged, and the static clip is formed, the images before and after conversion are compared to judge, the wrong clip is marked and an alarm is issued.
It realizes the rapid and accurate identification of video conversion errors and marks the error locations, which facilitates subsequent repairs.
Smart Images

Figure CN120264002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video image processing, and particularly to an image recognition processing method and system applied to video conversion. Background Art
[0002] Currently, the recognition and processing of images are two necessary steps after such intelligent systems receive images. Image recognition involves how a computer recognizes and classifies objects, scenes, and features in an image. Image recognition technology may be based on the main features of an image, including color, texture, shape, etc. Image processing involves operating on and analyzing an image to extract useful information or improve image quality. Common image processing operations include image enhancement, image compression, image segmentation, and image synthesis, etc. Video conversion usually refers to the conversion of video formats, usually to adapt to a player or playback device, or to perform format conversion according to user requirements.
[0003] The above-mentioned prior art solutions have the following defects: During the video conversion process, errors may occur, resulting in errors in some segments of the video, such as image garbled characters or video segment misalignment. It is difficult for the prior art to automatically identify the position where the video conversion goes wrong. Summary of the Invention
[0004] In order to conveniently and quickly identify whether an error occurs after video conversion and the position of the error, the present application provides an image recognition processing method and system applied to video conversion.
[0005] On the one hand, an image recognition processing method applied to video conversion provided by the present application adopts the following technical solution: An image recognition processing method applied to video conversion includes the following steps: Preset video slicing criteria; Receive the input video information, and intercept multiple images in the video as reference images according to the preset video slicing criteria; Judge whether each frame of image before and after the reference image in the video information is the same as the reference image, select the images that are the same as the reference image and can be connected to each other, combine the selected images corresponding to each reference image to form a video segment, and use the video segment as a static segment; If a static segment contains multiple reference images at the same time, retain the reference image closest to the start of the video; Perform format conversion on the video information; Intercept comparison segments from the converted video information according to the time periods corresponding to all static segments, and compare the comparison segments with the static segments; If all comparison segments are the same as the static segments in the same time period, it is judged that the video format conversion is successful; If any comparison segment is different from the static segment in the same time period, it is determined that an error has occurred in the video format conversion, an alarm is issued, and the different comparison segment and static segment are marked.
[0006] By adopting the above solution, after the video conversion, the system will determine whether an error has occurred in the video conversion based on the reference image before conversion and the sliced images after conversion. When an error is found, the error part will be marked and the corresponding part of the original video will be marked, which is convenient for subsequent video repair. Since the judgment of whether there is an error in the video conversion is based on the intercepted images, the images can be compared even after format conversion, so the error judgment of this system for video conversion is accurate and fast.
[0007] Preferably, the step of "presetting video slicing criteria" further includes: Presetting a time interval range; Setting the video slicing criteria as: within the time interval range, find the two frames with the largest change in the picture from the beginning of the video information, and select the image closest to the frame at the beginning of the video as the reference image.
[0008] By adopting the above solution, according to the video slicing criteria, a video segment will be first drawn from the beginning of the video according to the time interval, then the two frames with the largest change in the picture will be found from the video segment, and the image of the first frame will be used as the reference image, and then the video segment will be divided according to the time interval starting from the other frame. In this way, the change between one reference image and the next reference image can be ensured to be large enough, and in the process of generating static segments using the reference images later, it can be ensured that there is a relatively large difference between two adjacent static segments in the video. When verifying whether an error has occurred in the video conversion, adjacent static segments in the video are not easily affected by each other.
[0009] Preferably, it further includes the following steps: If it is determined that an error has occurred in the video format conversion, it is determined whether the marked comparison segment is a static video; If the marked comparison segment is a static video, intercept the picture of any frame in the comparison segment as the comparison picture, and find the closest reference image according to the comparison picture; Find the corresponding static segment according to the found reference image, judge the relative position of the time period of the found static segment and the time period where the comparison segment is located, and mark the found static segment; If the marked comparison segment is a dynamic video, intercept multiple pictures in the comparison segment as comparison pictures according to a predetermined rule, and find the closest reference image according to each comparison picture; Delete the reference images that have a large difference from the corresponding comparison images, find the corresponding static segments based on the remaining reference images, and determine the relative position of the time period of the found static segment and the time period where the comparison segment is located. If there is an overlapping part between the time period of the found static segment and the time period where the comparison segment is located, then delete the overlapping part of the comparison segment and mark it as an error segment.
[0010] By adopting the above solution, the system can intelligently determine whether the comparison segment is a segment of other parts in the original video information, that is, whether there is a dislocation in the video conversion, by comparing and judging whether it is a static video or a dynamic video, and prompt the possible dislocation information, which further facilitates video repair.
[0011] Preferably, the following steps are further included: Read the audio information in the received video information and convert it into text information, and add the text information to the video information in the time order of the audio information; Read the audio information in the converted video information and convert it into text information, and add the text information to the video information in the time order of the audio information; If an error occurs in the video format conversion, then compare the text information in the marked comparison segment and the static segment; If the text information in the marked comparison segment is the same as the text information in the static segment, then retain the audio information in the comparison segment, perform format conversion on the video segment before the two static segments connected to the marked static segment, and replace the part with the same time axis in the converted video information.
[0012] By adopting the above solution, when intelligent subtitles need to be generated in the video, the system can use the subtitles to compare the audio before and after the video conversion, and can more accurately retrieve whether there is an error in the video and the relevant information about the error.
[0013] Preferably, the following steps are further included: If the text information in the marked comparison segment is different from the text information in the static segment, then judge whether the language logic of the text information in the comparison segment is correct; If the language logic of the text information in the comparison segment is incorrect, then judge that the audio information in the comparison segment is incorrect; If the language logic of the text information in the comparison segment is correct, then search for a segment in the received video information that is the same as the text information in the comparison segment; If a segment that is the same as the text information in the comparison segment is found, then mark this segment as a repair reference segment.
[0014] By adopting the above solution, the system can make a rationality judgment on the text information through common language logic, and can effectively reduce the influence of errors generated during audio-to-text conversion on the system's judgment.
[0015] On the other hand, an image recognition processing system applied to video conversion provided by the present application adopts the following technical solutions: An image recognition processing system applied to video conversion includes a data storage module, a video slicing module, a video conversion module, a video detection module, an error detection module, and a video playback module; The data storage module prestores video slicing standards; The video slicing module receives the input video information, calls the video slicing standard of the data storage module, intercepts multiple images in the video as reference images according to the video slicing standard, determines whether the images of each frame before and after the reference image in the video information are the same as the reference image, selects the images that are the same as the reference image and can be connected to each other, combines the selected images corresponding to each reference image to form a video segment, takes the video segment as a static segment. If there are multiple reference images in a static segment, the reference image closest to the start of the video is retained, and the reference image and the static segment are transmitted to the data storage module for storage; The video conversion module receives the input video information and performs format conversion on the video information; The video detection module calls the video information before conversion and the video information after conversion of the video conversion module, calls the reference image and the static segment of the data storage module according to the video information before conversion, intercepts comparison segments from the video information after conversion according to the time periods corresponding to all the static segments, compares the comparison segments with the static segments. If all the comparison segments are the same as the static segments in the same time period, a success signal is transmitted to the video playback module. If any comparison segment is different from the static segment in the same time period, error video information is transmitted to the error detection module; After receiving the error video information, the error detection module marks the different comparison segments and static segments and transmits the marked video information to the video playback module; After receiving the success signal, the video playback module outputs the video information after conversion. After receiving the marked comparison segments and static segments, it issues an alarm and outputs the marked comparison segments and static segments.
[0016] By adopting the above solution, after the video conversion, the system will judge whether there is an error in the video conversion according to the reference image before conversion and the sliced images after conversion. When an error is found, the error part will be marked and the corresponding part of the original video will also be marked, which is convenient for subsequent video repair. Since the judgment of whether there is an error in the video conversion is based on the intercepted images, the images can be compared even after format conversion. Therefore, the error judgment of the video conversion by this system is accurate and fast.
[0017] Preferably, the data storage module is preset with a time interval range and receives the input video slice standard, where the video slice standard is: within the time interval range, find the two frames with the largest change in the picture from the beginning of the video information, and select the image closest to the frame at the beginning of the video as the reference image.
[0018] By adopting the above solution, according to the video slice standard, a video segment will be first delimited from the beginning of the video according to the time interval, then find the two frames with the largest change in the picture from the video segment, and use the image of the first frame as the reference image, and then delimit the video segment according to the time interval starting from another frame. In this way, the change between one reference image and the next reference image can be ensured to be large enough, and it can be ensured that there is a relatively large difference between two adjacent static segments in the video during the subsequent process of generating static segments using the reference image. When verifying whether there is an error in the video conversion, the adjacent static segments in the video are not easily affected by each other.
[0019] Preferably, it further includes a video repair module; The error detection module determines whether the marked comparison segment is a static video. If the marked comparison segment is a static video, intercept any frame of the comparison segment as the comparison picture, find the closest reference image according to the comparison picture, and transmit the reference image and the static signal to the video repair module; if the marked comparison segment is a dynamic video, intercept multiple pictures of the comparison segment as the comparison pictures according to a predetermined rule, find the closest reference image according to each comparison picture, and transmit the reference image and the dynamic signal to the video repair module; After receiving the reference image and the static signal, the video repair module finds the corresponding static segment according to the found reference image, determines the relative position of the time period of the found static segment and the time period where the comparison segment is located, and marks the found static segment; after receiving the reference image and the dynamic signal, the video repair module deletes the reference images with large differences from the corresponding comparison pictures, finds the corresponding static segments according to the remaining reference images, determines the relative position of the time period of the found static segment and the time period where the comparison segment is located. If there is an overlapping part between the time period of the found static segment and the time period where the comparison segment is located, then delete the overlapping part of the comparison segment and mark it as an error segment.
[0020] By adopting the above solution, the system can intelligently determine whether the comparison segment is a segment of other parts in the original video information, that is, whether there is a misalignment in the video conversion, by comparing whether it is a static video or a dynamic video, and prompt the possible misalignment information, which further facilitates video repair.
[0021] Preferably, it further includes a text generation module. The text generation module calls the video information received by the video slicing module and the converted video information of the video conversion module, reads the audio information in the received video information and converts it into text information, adds the text information to the video information in the time sequence of the audio information and transmits it to the video slicing module, reads the audio information in the converted video information and converts it into text information, adds the text information to the video information in the time sequence of the audio information and transmits it to the video repair module; The video repair module compares the text information in the marked comparison segment and the static segment. If the text information in the marked comparison segment is the same as the text information in the static segment, the audio information in the comparison segment is retained, and the video segment before the two static segments connected to the marked static segment is format-converted and replaces the part with the same time axis in the converted video information.
[0022] By adopting the above solution, when intelligent subtitles need to be generated in the video, the system can use the subtitles to compare the audio before and after video conversion, and can more accurately retrieve whether there are errors in the video and the relevant information about the errors.
[0023] Preferably, if the text information in the marked comparison segment is different from the text information in the static segment, the video repair module determines whether the language logic of the text information in the comparison segment is correct. If the language logic of the text information in the comparison segment is incorrect, it determines that the audio information in the comparison segment is incorrect. If the language logic of the text information in the comparison segment is correct, it searches for a segment in the received video information that is the same as the text information in the comparison segment. If a segment that is the same as the text information in the comparison segment is found, the segment is marked as a repair reference segment.
[0024] By adopting the above solution, the system can judge the rationality of the text information through common language logic, and can effectively reduce the influence of errors generated during audio-to-text conversion on the system judgment.
[0025] In summary, the present invention has the following beneficial effects: 1. After the video is converted, the system will judge whether there is an error in the video conversion according to the reference image before conversion and the sliced image after conversion. When an error is found, the error part will be marked and the corresponding part of the original video will also be marked, which is convenient for subsequent video repair. Since the judgment of whether there is an error in the video conversion is based on the intercepted images, the images can be compared even after format conversion, so the error judgment of the video conversion by this system is accurate and fast. Description of the Drawings
[0026] Figure 1 It is the overall flowchart of Embodiment 1 of this application.
[0027] Figure 2 It is the overall system block diagram of the second embodiment of this application.
[0028] Explanation of reference numerals: 1. Data storage module; 2. Video slicing module; 3. Video conversion module; 4. Video detection module; 5. Error detection module; 6. Video playback module; 7. Video repair module; 8. Text generation module. Specific implementation manners
[0029] Embodiment 1. An image recognition processing method applied to video conversion disclosed in an embodiment of this application is as Figure 1 shown, and the specific steps are as follows: S100. Requirement presetting: S101. Preset the time interval range. Set the video slicing standard as: within the time interval range, start searching for the two frames with the largest change in the video information from the beginning of the video, and select the image closest to the frame at the beginning of the video as the reference image. According to the video slicing standard, first divide a video segment from the beginning of the video according to the time interval, then find the two frames with the largest change in the video segment, and use the image of the first frame as the reference image, and then divide the video segment according to the time interval starting from the other frame. In this way, the change between one reference image and the next reference image can be ensured to be large enough, and it can be ensured that there is a relatively large difference between two adjacent static segments in the video during the process of generating static segments using the reference image. When verifying whether there is an error in the video conversion, adjacent static segments in the video are not easily affected by each other.
[0030] S200. Video preprocessing: S201. Receive the input video information, and intercept multiple images in the video as reference images according to the preset video slicing standard.
[0031] S202. Judge whether each frame of the image before and after the reference image in the video information is the same as the reference image, select the images that are the same as the reference image and can be connected to each other, combine the selected images corresponding to each reference image to form a video segment, and use the video segment as a static segment.
[0032] S203. If a static segment contains multiple reference images at the same time, retain the reference image closest to the beginning of the video.
[0033] S204. Read the audio information in the received video information and convert it into text information, and add the text information to the video information in the time sequence of the audio information.
[0034] S300. Format conversion: S301. Perform format conversion on the video information.
[0035] S400, text generation: S401, reading the audio information in the converted video information and converting it into text information, and adding the text information into the video information according to the time sequence of the audio information.
[0036] S500, conversion quality judgment: S501: Extract comparison segments from the converted video information according to time periods corresponding to all static segments, and compare the comparison segments with the static segments.
[0037] S600, correct judgment: S601: If all the compared segments are the same as the static segments in the same time period, it is determined that the video format conversion is successful.
[0038] S700, judgment error: S701: If any comparison segment is different from the static segment in the same time period, it is determined that an error occurs in the video format conversion, an alarm is issued, and the different comparison segments and static segments are marked.
[0039] S800, Error Diagnosis: S801: Determine whether the marked comparison segment is a static video.
[0040] S802: If the marked comparison segment is a static video, capture any frame in the comparison segment as a comparison picture, and search for the closest reference image based on the comparison picture.
[0041] S803: Search for a corresponding static segment according to the found reference image, determine the relative position of the time period of the found static segment and the time period of the comparison segment, and mark the found static segment.
[0042] S804: If the marked comparison segment is a dynamic video, multiple frames in the comparison segment are intercepted as comparison frames according to a predetermined rule, and the closest reference image is searched for each comparison frame.
[0043] S805. Delete the reference image that is greatly different from the corresponding comparison image, search for the corresponding static segment based on the remaining reference image, and determine the relative position of the time period of the static segment found and the time period of the comparison segment. If the time period of the static segment found overlaps with the time period of the comparison segment, delete the overlapping part of the comparison segment and mark it as an error segment.
[0044] S900, text-assisted comparison: S901: If an error occurs in the video format conversion, compare the text information in the marked comparison segment and the static segment.
[0045] S902. If the text information in the marked comparison segment is the same as the text information in the static segment, retain the audio information in the comparison segment, perform format conversion on the video segments before the two static segments connected to the marked static segment, and replace the parts with the same time axis in the converted video information. When intelligent subtitles need to be generated in the video, the system can compare the audio before and after the video conversion using the subtitles, which can more accurately retrieve whether there are errors in the video and the relevant information about the errors.
[0046] S903. If the text information in the marked comparison segment is different from the text information in the static segment, determine whether the language logic of the text information in the comparison segment is correct.
[0047] S904. If the language logic of the text information in the comparison segment is incorrect, determine that the audio information in the comparison segment is incorrect.
[0048] S905. If the language logic of the text information in the comparison segment is correct, search for segments with the same text information as the text information in the comparison segment in the received video information.
[0049] S906. If a segment with the same text information as the text information in the comparison segment is found, mark this segment as a repair reference segment. The system can make a rationality judgment on the text information through common language logic, which can effectively reduce the influence of errors generated during audio-to-text conversion on the system's judgment.
[0050] The implementation principle of an image recognition processing method for video conversion in an embodiment of this application is as follows: After video conversion, the system will determine whether there are errors in the video conversion based on the reference image before conversion and the sliced image after conversion. When errors are found, the error parts will be marked and the corresponding parts of the original video will also be marked, which is convenient for subsequent video repair. Since the judgment of whether there are errors in video conversion is based on the intercepted images, the images can be compared even after format conversion, so the error judgment of this system for video conversion is accurate and fast.
[0051] Embodiment 2. An embodiment of this application discloses an image recognition processing system for video conversion, as Figure 2 shown, including a data storage module 1, a video slicing module 2, a video conversion module 3, a video detection module 4, an error detection module 5, a video playback module 6, a video repair module 7, and a text generation module 8.
[0052] The data storage module 1 pre-stores a time interval range and receives the input video slicing standard. The video slicing standard is as follows: within the time interval range, find the two frames with the largest change in the picture from the beginning of the video information, and select the image closest to the beginning frame of the video as the reference image. According to the video slicing standard, first, a video segment is divided from the beginning of the video according to the time interval, then find the two frames with the largest change in the picture from the video segment, and use the image of the first frame as the reference image. Then, start from another frame and divide the video segment according to the time interval. In this way, the change between one reference image and the next reference image can be ensured to be large enough, so that there is a relatively large difference between two adjacent static segments in the video during the subsequent process of generating static segments using the reference image. When verifying whether there is an error in video conversion, adjacent static segments in the video are not easily affected by each other.
[0053] The video slicing module 2 receives the input video information, calls the video slicing standard of the data storage module 1, intercepts multiple images in the video as reference images according to the video slicing standard, determines whether each frame of the image before and after the reference image in the video information is the same as the reference image, selects the images that are the same as the reference image and can be connected to each other, combines the selected images corresponding to each reference image to form a video segment, and uses the video segment as a static segment. If a static segment contains multiple reference images at the same time, the same reference images are combined into one reference image. If a static segment contains multiple reference images at the same time, the static segment corresponding to the reference image closest to the beginning of the video is also retained and combined. The reference image and the static segment are transmitted to the data storage module 1 for storage.
[0054] The video conversion module 3 receives the input video information and performs format conversion on the video information.
[0055] The text generation module 8 calls the video information received by the video slicing module 2 and the video information after conversion by the video conversion module 3, reads the audio information in the received video information and converts it into text information, adds the text information to the video information in the time sequence of the audio information and transmits it to the video slicing module 2, reads the audio information in the converted video information and converts it into text information, adds the text information to the video information in the time sequence of the audio information and transmits it to the video repair module 7.
[0056] The video detection module 4 calls the video information before conversion and after conversion of the video conversion module 3, calls the reference image and static segments of the data storage module 1 according to the video information before conversion, intercepts comparison segments from the video information after conversion according to the time periods corresponding to all static segments, compares the comparison segments with the static segments. If all comparison segments are the same as the static segments in the same time period, it transmits a success signal to the video playback module 6. If any comparison segment is different from the static segment in the same time period, it transmits error video information to the error detection module 5.
[0057] After receiving the error video information, the error detection module 5 marks the different comparison segments and static segments and transmits the marked video information to the video playback module 6. The error detection module 5 determines whether the marked comparison segment is a static video. If the marked comparison segment is a static video, it intercepts the picture of any frame in the comparison segment as the comparison picture, searches for the closest reference image according to the comparison picture, and transmits the reference image and the static signal to the video repair module 7. If the marked comparison segment is a dynamic video, it intercepts multiple pictures in the comparison segment as the comparison pictures according to a predetermined rule, searches for the closest reference image according to each comparison picture, and transmits the reference image and the dynamic signal to the video repair module 7.
[0058] After receiving the success signal, the video playback module 6 outputs the video information after conversion. After receiving the marked comparison segments and static segments, it issues an alarm and outputs the marked comparison segments and static segments.
[0059] After receiving the reference image and the static signal, the video repair module 7 searches for the corresponding static segment according to the found reference image, judges the relative position of the time period of the found static segment and the time period where the comparison segment is located, and marks the found static segment. After receiving the reference image and the dynamic signal, the video repair module 7 deletes the reference images with large differences from the corresponding comparison pictures, searches for the corresponding static segments according to the remaining reference images, judges the relative position of the time period of the found static segment and the time period where the comparison segment is located. If there is an overlapping part between the time period of the found static segment and the time period where the comparison segment is located, it deletes the overlapping part of the comparison segment and marks it as an error segment. The system can intelligently judge whether the comparison segment is a segment of other parts in the original video information, that is, whether there is a dislocation in the video conversion, by comparing and judging whether it is a static video or a dynamic video, and prompts the possible dislocation information, which further facilitates video repair.
[0060] The video repair module 7 compares the text information in the marked comparison segment and the static segment. If the text information in the marked comparison segment is the same as the text information in the static segment, the audio information in the comparison segment is retained, and the video segment before the two static segments connected to the marked static segment is format-converted and the part with the same time axis in the converted video information is replaced. If the text information in the marked comparison segment is different from the text information in the static segment, the video repair module 7 determines whether the language logic of the text information in the comparison segment is correct. If the language logic of the text information in the comparison segment is incorrect, it is determined that the audio information in the comparison segment is incorrect. If the language logic of the text information in the comparison segment is correct, a segment with the same text information as that in the comparison segment is searched for in the received video information. If a segment with the same text information as that in the comparison segment is found, that segment is marked as a repair reference segment.
[0061] The embodiments of the specific implementation manners are all preferred embodiments of the present invention, and do not limit the protection scope of the present invention accordingly. Therefore, all equivalent changes made according to the structure, shape, and principle of the present invention shall be covered by the protection scope of the present invention.
Claims
1. An image recognition processing method applied to video conversion, characterized in that, Including the following steps: Presetting a video slicing standard; Receiving the input video information, and intercepting multiple images in the video as reference images according to the preset video slicing standard; Judging whether each frame of the image before and after the reference image in the video information is the same as the reference image, selecting the images that are the same as the reference image and can be connected to each other, combining the selected images corresponding to each reference image to form a video segment, and taking the video segment as a static segment; If a static segment contains multiple reference images at the same time, retain the reference image closest to the start of the video; Performing format conversion on the video information; Intercepting comparison segments according to the time periods corresponding to all the static segments for the converted video information, and comparing the comparison segments with the static segments; If all the comparison segments are the same as the static segments in the same time period, it is judged that the video format conversion is successful; If any comparison segment is different from the static segment in the same time period, it is judged that an error occurs in the video format conversion, an alarm is issued, and the different comparison segments and static segments are marked.
2. The image recognition processing method and system for video conversion according to claim 1, wherein, The step of "presetting a video slicing standard" further includes: Presetting a time interval range; Setting the video slicing standard as: searching for two frames with the largest change in the picture from the beginning of the video information within the time interval range, and selecting the image closest to the frame at the start of the video as the reference image.
3. An image recognition processing method applied to video conversion according to claim 1, characterized in that, Further including the following steps: If it is judged that an error occurs in the video format conversion, judging whether the marked comparison segment is a static video; If the marked comparison segment is a static video, intercepting a frame of the picture in the comparison segment as a comparison picture, and searching for the closest reference image according to the comparison picture; Searching for the corresponding static segment according to the found reference image, judging the relative position of the time period of the found static segment and the time period where the comparison segment is located, and marking the found static segment; If the marked comparison segment is a dynamic video, intercepting multiple pictures in the comparison segment as comparison pictures according to a predetermined rule, and searching for the closest reference image according to each comparison picture; Deleting the reference images with large differences from the corresponding comparison pictures, searching for the corresponding static segments according to the remaining reference images, judging the relative position of the time period of the found static segment and the time period where the comparison segment is located, and if there is an overlapping part between the time period of the found static segment and the time period where the comparison segment is located, deleting the overlapping part of the comparison segment and marking it as an error segment.
4. An image recognition processing method applied to video conversion according to claim 1, characterized in that, Further including the following steps: Reading the audio information in the received video information and converting it into text information, and adding the text information to the video information in the time order of the audio information; Reading the audio information in the converted video information and converting it into text information, and adding the text information to the video information in the time order of the audio information; If an error occurs in the video format conversion, comparing the text information in the marked comparison segment and the static segment; If the text information in the marked comparison segment is the same as the text information in the static segment, retaining the audio information in the comparison segment, and performing format conversion on the video segment before the two static segments connected to the marked static segment and replacing the part with the same time axis in the converted video information.
5. An image recognition processing method applied to video conversion according to claim 4, characterized in that, Further including the following steps: If the text information in the marked comparison segment is different from the text information in the static segment, then determine whether the language logic of the text information in the comparison segment is correct; If the language logic of the text information in the comparison segment is incorrect, then determine that the audio information in the comparison segment is incorrect; If the language logic of the text information in the comparison segment is correct, then search for segments in the received video information that are the same as the text information in the comparison segment; If a segment that is the same as the text information in the comparison segment is found, then mark this segment as a repair reference segment.
6. An image recognition processing system applied to video conversion, characterized in that: It includes a data storage module (1), a video slicing module (2), a video conversion module (3), a video detection module (4), an error detection module (5), and a video playback module (6); The data storage module (1) prestores video slicing criteria; The video slicing module (2) receives the input video information, calls the video slicing criteria of the data storage module (1), intercepts multiple images in the video as reference images according to the video slicing criteria, determines whether each frame of the image before and after the reference image in the video information is the same as the reference image, selects the images that are the same as the reference image and can be connected to each other, combines the selected images corresponding to each reference image to form a video segment, uses the video segment as a static segment. If there are multiple reference images in one static segment, then retain the reference image closest to the start of the video, and transmit the reference image and the static segment to the data storage module (1) for storage; The video conversion module (3) receives the input video information and performs format conversion on the video information; The video detection module (4) calls the video information before conversion and the video information after conversion of the video conversion module (3), calls the reference image and the static segment of the data storage module (1) according to the video information before conversion, intercepts comparison segments from the video information after conversion according to the time periods corresponding to all the static segments, compares the comparison segments with the static segments. If all the comparison segments are the same as the static segments in the same time period, then transmit a success signal to the video playback module (6). If any comparison segment is different from the static segment in the same time period, then transmit the error video information to the error detection module (5); After receiving the error video information, the error detection module (5) marks the different comparison segments and static segments and transmits the marked video information to the video playback module (6); After receiving the success signal, the video playback module (6) outputs the video information after conversion. After receiving the marked comparison segments and static segments, it issues an alarm and outputs the marked comparison segments and static segments.
7. An image recognition processing system applied to video conversion according to claim 6, characterized in that: The data storage module (1) presets a time interval range and receives the input video slicing criteria. The video slicing criteria are: within the time interval range, search for the two frames with the largest change in the picture from the start of the video information, and select the image closest to the frame at the start of the video as the reference image.
8. An image recognition processing system applied to video conversion according to claim 6, characterized in that: It also includes a video repair module (7); The error detection module (5) determines whether the marked comparison segment is a static video. If the marked comparison segment is a static video, it intercepts a frame of the comparison segment as the comparison picture, searches for the closest reference image according to the comparison picture, and transmits the reference image and the static signal to the video repair module (7); if the marked comparison segment is a dynamic video, it intercepts multiple pictures in the comparison segment as the comparison pictures according to a predetermined rule, searches for the closest reference image according to each comparison picture, and transmits the reference image and the dynamic signal to the video repair module (7); After receiving the reference image and the static signal, the video repair module (7) searches for the corresponding static segment according to the found reference image, determines the relative position of the time period of the found static segment and the time period where the comparison segment is located, and marks the found static segment; after receiving the reference image and the dynamic signal, the video repair module (7) deletes the reference images with large differences from the corresponding comparison pictures, searches for the corresponding static segments according to the remaining reference images, determines the relative position of the time period of the found static segment and the time period where the comparison segment is located. If there is an overlapping part between the time period of the found static segment and the time period where the comparison segment is located, the overlapping part of the comparison segment is deleted and marked as an error segment.
9. An image recognition processing system applied to video conversion according to claim 8, characterized in that: It further includes a text generation module (8). The text generation module (8) calls the video information received by the video slicing module (2) and the converted video information of the video conversion module (3), reads the audio information in the received video information and converts it into text information, adds the text information to the video information in the time sequence of the audio information and transmits it to the video slicing module (2), reads the audio information in the converted video information and converts it into text information, adds the text information to the video information in the time sequence of the audio information and transmits it to the video repair module (7); The video repair module (7) compares the text information in the marked comparison segment and the static segment. If the text information in the marked comparison segment is the same as the text information in the static segment, it retains the audio information in the comparison segment, performs format conversion on the video segment before the two static segments connected to the marked static segment and replaces the part with the same time axis in the converted video information.
10. An image recognition processing system applied to video conversion according to claim 9, characterized in that: If the text information in the marked comparison segment is different from the text information in the static segment, the video repair module (7) determines whether the language logic of the text information in the comparison segment is correct. If the language logic of the text information in the comparison segment is incorrect, it determines that the audio information in the comparison segment is incorrect. If the language logic of the text information in the comparison segment is correct, it searches for a segment with the same text information as the text information in the comparison segment in the received video information. If a segment with the same text information as the text information in the comparison segment is found, the segment is marked as a repair reference segment.