Character generation method based on image recognition
By preprocessing and reprocessing image and voice information, combined with image enhancement and voice segmentation technology, the accuracy and scene adaptability problems of image recognition text generation in existing technologies are solved, achieving higher text generation accuracy and user experience.
Patent Information
- Application Number
- CN202510749883.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-09
AI Technical Summary
Existing text generation methods based on image recognition lack accuracy in terms of image quality and environmental cross-validation, resulting in inaccurate generated text content and incompatibility with the scene.
By obtaining the original image and environmental voice information of the area to be collected, preprocessing and reprocessing are performed, and combining image enhancement and voice cutting technology, the text generation results are adjusted to match the scene type to ensure the accuracy and adaptability of the generated text.
The accuracy and scene adaptability of image recognition text generation have been improved, providing a good user viewing experience.
Smart Images

Figure CN120612461A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular to a text generation method based on image recognition. Background Art
[0002] Currently, image recognition refers to the use of computers to process, analyze, and understand images to identify various patterns of targets and objects. It is a practical application of deep learning algorithms. Text generation refers to the generation of text content by artificial intelligence. This is achieved by using various machine learning methods to learn the components of objects from data and then generate new and completely original text content.
[0003] The text generation method based on image recognition in the existing technology refers to image preprocessing after image acquisition, and then detecting, locating and identifying the text, and performing post-processing and semantic enhancement. Among them, text recognition uses features such as the shape, edge texture, etc. of the characters for recognition. This method improves the recognition accuracy to a certain extent, but the accuracy of feature extraction is still limited by the image quality. It does not take into account whether the text content generated by environmental cross-validation is accurate, and does not take into account whether the scene of the generated text is suitable. The accuracy of text generation based on image recognition is low, and there is room for improvement. Summary of the Invention
[0004] In order to improve the accuracy of text generation based on image recognition, the present application provides a text generation method based on image recognition.
[0005] The text generation method based on image recognition provided in this application adopts the following technical solutions:
[0006] The text generation method based on image recognition includes the following steps:
[0007] Step S1, obtaining a to-be-collected area, determining a related environmental area based on the to-be-collected area, collecting information from the to-be-collected area and the related environmental area to obtain original collected information, wherein the original collected information includes original image information and original environmental voice information;
[0008] Step S2, preprocessing the original image information to obtain image preprocessing information, and preprocessing the original environmental voice information to obtain environmental voice preprocessing information;
[0009] Step S3, reprocessing the preprocessed image information to obtain reprocessed image information of the acquisition area, and generating a preliminary text generation result based on the reprocessed image information of the acquisition area;
[0010] Step S4, judging whether the preliminary text generation result needs to be adjusted based on the image preprocessing information and the environmental voice preprocessing information, and if so, performing the adjustment to obtain a text generation environment adjustment result;
[0011] Step S5, analyzing and identifying the scene when collecting information in the collection area based on the original image information and the original environmental voice information to obtain collection scene identification information;
[0012] Step S6, based on the acquisition scene recognition information, determines whether the scene during image acquisition matches the scene of the text generation environment adjustment result, and determines whether adjustment is needed. If adjustment is needed, adjust the text playback speed and image playback speed until the scenes match, and output the final text generation result.
[0013] Preferably, the relevant environmental area is determined based on the area of the area to be collected and the scene complexity;
[0014] The original image information is obtained by performing image acquisition operation based on a high-definition camera, wherein the original image information includes image information of the acquisition area and image information of the related environmental area;
[0015] When performing an image acquisition operation, the ambient sound around the acquisition point is synchronously acquired to obtain original ambient voice information, and the original image information at the same time is matched with the original ambient voice information based on the acquisition time.
[0016] Preferably, a preprocessing operation is performed on the acquisition area image information and the related environment area image information in the original image information based on an image noise reduction technology to obtain image preprocessing information, wherein the image preprocessing information includes the acquisition area preprocessing image information and the related environment preprocessing image information;
[0017] Based on the speech segmentation technology, the original environmental speech information is segmented to obtain a speech logical segment dataset, where the speech logical segment dataset includes multiple speech logical segments. Based on the time of each speech logical segment, whether the speech logical segments overlap is determined. If overlap, the timbre and content fluency of each speech logical segment are used to determine whether it is mis-segmented. If it is mis-segmented, the speech logical segment dataset is adjusted and updated;
[0018] Determine whether there is human voice in each voice logic segment, filter out the voice logic segments without human voice and delete them, filter out the voice logic segments with human voice and mark them as valid voice logic segments, and pre-process the valid voice logic segments based on pre-emphasis processing technology and windowing processing technology to obtain environmental voice preprocessing information.
[0019] Preferably, based on image enhancement technology and image distortion correction algorithm, the pre-processed image information of the acquisition area is further processed to obtain re-processed image information of the acquisition area;
[0020] The image information is processed based on the collected area to initially generate the preliminary text generation result.
[0021] Preferably, the area containing text in the collected area reprocessed image information is marked as a text content area, and the text content of each text content area is subjected to feature extraction to obtain a plurality of text feature extraction information;
[0022] Based on the distribution and arrangement of each text content area in the area to be collected and the font, the importance of the text content in each text content area is determined to obtain the text generation priority of each text content area;
[0023] Based on the text generation priority of each text content area, preliminary text generation is performed on the text feature extraction information of each text content area in turn to obtain a preliminary text generation result.
[0024] Preferably, a first type of adjustment is performed on the preliminary text generation result according to the relevant environmental image preprocessing information to obtain a first type of text generation adjustment result;
[0025] Performing a second type of adjustment on the first type of adjusted text generation result according to the environmental speech preprocessing information to obtain a second type of text generation adjustment result;
[0026] Among them, the first type of text generation adjustment results and the second type of text generation adjustment results are combined to form the text generation environment adjustment results.
[0027] Preferably, a scene type data pool is obtained based on big data technology, wherein the scene type data pool includes user original type, brand marketing type, knowledge popularization type, news information type, and entertainment interaction type;
[0028] The original image information and the original environmental voice information are matched with the scene type data pool to determine the scene when the information is collected in the area to be collected to obtain the collection scene identification information.
[0029] Preferably, the text generation environment adjustment result is matched with the scene type data pool to determine the scene to which the generated text is suitable to obtain text generation scene identification information;
[0030] Match the text generation scene recognition information with the collection scene recognition information. If the text generation scene recognition information matches the collection scene recognition information, there is no need to adjust the output of the final text generation result. If the text generation scene recognition information does not match the collection scene recognition information, adjust the generated text playback speed. When the generated text playback speed is adjusted to the preset text playback speed threshold, stop adjusting the generated text playback speed and start adjusting the image playback speed until the text generation scene recognition information matches the collection scene recognition information during playback, and output the final text generation result.
[0031] In summary, this application includes at least one of the following beneficial technical effects:
[0032] 1. Preprocessing the original image information through image noise reduction to obtain image preprocessing information improves image quality, thereby improving the accuracy of subsequent text generation based on image recognition. The original ambient voice information is segmented through voice segmentation technology, and it is determined whether it is mis-segmented. If it is mis-segmented, adjustments are made to improve voice quality. By determining whether there is human voice in the voice logical segment, the voice logical segment containing human voice is marked as a valid voice logical segment. The valid voice logical segment is preprocessed based on pre-emphasis processing technology and windowing technology to obtain ambient voice preprocessing information, further improving the ambient voice quality.
[0033] 2. Using relevant environmental image preprocessing information, the initial text generation results are adjusted to obtain a first type of adjustment result. The images of the area to be collected and the images of the relevant environmental areas are used to determine whether the initial text generation results contain errors. If errors occur, adjustments are made, thereby improving the accuracy of text generation based on image recognition. The first type of text generation adjustment results are then adjusted using environmental speech preprocessing information to obtain a second type of text generation adjustment result. The environmental speech during image collection is used to determine whether the first type of text generation adjustment results contain errors. If errors occur, adjustments are made, further improving the accuracy of text generation based on image recognition.
[0034] 3. Obtain text generation scene recognition information from the scene corresponding to the text generated by the text generation scene recognition information, match the collected scene recognition information with the text generation scene recognition information, and determine that when the image of the area to be collected is played, the scene corresponding to the collected information and the scene function corresponding to the generated text are consistent, so as to facilitate the accurate matching of the generated text with the corresponding image and voice, and facilitate the user's good viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a flow chart of a method for generating text based on image recognition, which is mainly reflected in this embodiment;
[0036] Figure 2 This is a schematic diagram of the process of step S32 in this embodiment;
[0037] Figure 3 This is a schematic diagram of the process of step S4 in this embodiment;
[0038] Figure 4 This is a flow chart mainly showing step S6 in this embodiment. DETAILED DESCRIPTION
[0039] The present application is further described in detail below with reference to the accompanying drawings.
[0040] The embodiment of the present application discloses a text generation method based on image recognition.
[0041] A method for generating text based on image recognition, comprising the following steps:
[0042] Reference Figure 1 In step S1, the area to be collected is obtained, and the related environmental area is determined based on the area to be collected. The information of the area to be collected and the related environmental area is collected to obtain the original collected information, which includes the original image information and the original environmental voice information. Step S1 specifically includes:
[0043] Step S11: Acquire the area to be collected, and determine the relevant environmental area based on the area of the area to be collected and the scene complexity.
[0044] Specifically, the area of the region to be collected is positively correlated with the area of the relevant environment, and the scene complexity of the region to be collected is positively correlated with the area of the relevant environment.
[0045] Among them, the scene complexity of the area to be collected is determined by the number of objects, the degree of difference in appearance between objects, the neatness of the object placement, and the brightness of light and shadow.
[0046] Based on the number of objects in the area to be collected, the influence of the number of objects on the scene complexity of the area to be collected is judged to obtain an object number correlation index, wherein the more objects in the area to be collected, the larger the object number correlation index.
[0047] Based on the degree of appearance difference between objects in the area to be collected, the impact of the object appearance difference on the scene complexity of the area to be collected is judged to obtain the object appearance difference correlation index, wherein the greater the degree of appearance difference between objects in the area to be collected, the greater the object appearance difference correlation index.
[0048] Based on the degree of neatness of the objects in the area to be collected, the influence of the object placement on the scene complexity of the area to be collected is judged to obtain the object placement related index, wherein the lower the degree of neatness of the objects in the area to be collected, the greater the object placement related index.
[0049] The maximum brightness difference is calculated according to the light and shadow brightness of each position in the area to be collected. Based on the maximum brightness, the influence of different light and shadow changes in the area to be collected on the scene complexity of the area to be collected is judged to obtain the light and shadow correlation index. Among them, the larger the maximum brightness difference, the larger the light and shadow correlation index.
[0050] The scene complexity of the area to be collected is judged based on the object quantity correlation index, object appearance difference correlation index, object placement correlation index and light and shadow correlation index. Among them, the object quantity correlation index, object appearance difference correlation index, object placement correlation index and light and shadow correlation index are all positively correlated with the scene complexity of the area to be collected.
[0051] Step S12: performing image acquisition based on a high-definition camera to obtain original image information, the original image information including acquisition area image information and related environment area image information. The acquisition area image information refers to the image within the area to be acquired, and the related environment area image information refers to the image of the related environment area.
[0052] Step S13: When performing the image acquisition operation, the ambient sound around the acquisition point is synchronously acquired to obtain the original ambient voice information, and the original image information at the same time is matched with the original ambient voice information based on the acquisition time.
[0053] In actual application, the relevant environmental area is determined by the area of the area with the collection area and the scene complexity. The relevant environmental area provides data support for subsequent information collection. Among them, the scene complexity of the area to be collected is determined by the number of objects, the degree of appearance difference between objects, the neatness of the object placement, and the brightness of light and shadow. This improves the accuracy of detecting and judging the scene complexity of the area to be collected, and thus improves the accuracy of the relevant environmental area. The original image information is obtained by image acquisition through a high-definition camera. At the same time, the ambient sound around the collection point is synchronously collected to obtain the original environmental voice information. The original image information and the original environmental voice information at the same time are matched at the same time, which improves the matching degree between the image and the voice, completes the image acquisition and voice acquisition, and provides data support for the subsequent text generation method based on image recognition.
[0054] Reference Figure 1 In step S2, the original image information is preprocessed to obtain image preprocessing information, and the original environmental voice information is preprocessed to obtain environmental voice preprocessing information. Step S2 specifically includes:
[0055] Step S21: Preprocess the image information of the acquisition area and the image information of the relevant environment area in the original image information using image noise reduction technology to obtain image preprocessing information, wherein the image preprocessing information includes the preprocessed image information of the acquisition area and the preprocessed image information of the relevant environment. Specifically, the image noise reduction technology can use Gaussian filtering for noise reduction.
[0056] Step S22: Based on the speech cutting technology, the original environmental speech information is cut to obtain a speech logic segment data set, wherein the speech logic segment data set includes multiple speech logic segments, and whether the speech logic segments overlap is determined based on the time of each speech logic segment. If they overlap, whether they are mis-segmented is determined based on the timbre and content fluency of each speech logic segment. If they are mis-segmented, the speech logic segment data set is adjusted and updated.
[0057] Let's take an example to illustrate: if the timbre of the sounds in the same speech logic segment is inconsistent, it is determined to be a mis-segmentation. The speech logic segment is segmented again based on the timbre in the speech logic segment, and the speech logic segment dataset is updated.
[0058] If the speech content in a speech logical segment is not fluent, that is, the pause interval of the speech content in the speech logical segment is greater than the preset speech pause interval threshold, the speech logical segment is re-segmented based on the speech pauses in the speech logical segment, and the speech logical segment dataset is updated.
[0059] Step S23, determine whether there is human voice in each voice logic segment, filter out the voice logic segments without human voice and delete them, filter out the voice logic segments with human voice and mark them as valid voice logic segments, and pre-process the valid voice logic segments based on pre-emphasis processing technology and windowing processing technology to obtain environmental voice preprocessing information.
[0060] In actual application, the original image information is preprocessed through image noise reduction processing to obtain image preprocessing information, which improves the image quality and thus improves the accuracy of subsequent text generation based on image recognition. The original environmental voice information is segmented through voice segmentation technology, and it is judged whether it is mis-segmented. If it is mis-segmented, adjustments are made to improve the voice quality. By judging whether there is human voice in the voice logic segment, the voice logic segment with human voice is marked as a valid voice logic segment, and the valid voice logic segment is preprocessed based on the pre-emphasis processing technology and windowing technology to obtain the environmental voice preprocessing information, which further improves the environmental voice quality.
[0061] Reference Figure 1In step S3, the pre-processed image information is processed again to obtain the re-processed image information of the acquisition area, and the preliminary text generation result is generated based on the re-processed image information of the acquisition area. Step S3 specifically includes:
[0062] Step S31 : Based on the image enhancement technology and the image distortion correction algorithm, the pre-processed image information of the acquisition area is further processed to obtain the re-processed image information of the acquisition area.
[0063] Specifically, image enhancement technology can use histogram equalization to enhance contrast, thereby highlighting the outline of text in the image and improving the accuracy of text generation based on image recognition.
[0064] In the embodiment of the present application, an image distortion correction algorithm is used to correct the pre-processed image information of the acquisition area where camera distortion may exist, thereby correcting the image deformation caused by the shooting angle, improving the image quality, and thereby improving the accuracy of text generation based on image recognition.
[0065] Step S32: reprocessing the image information based on the captured area to preliminarily generate text content to obtain a preliminary text generation result.
[0066] Wherein, step S32 specifically includes the following sub-steps:
[0067] Reference Figure 1 and Figure 2 In step S321, the area containing text in the collected area and the processed image information is marked as a text content area, and the text content of each text content area is subjected to feature extraction to obtain a plurality of text feature extraction information.
[0068] Step S322 : Based on the distribution and font of each text content area in the area to be collected, the importance of the text content in each text content area is determined to obtain the text generation priority of each text content area.
[0069] Specifically, if the text content areas are neatly distributed in the area to be collected, such as from top to bottom and from left to right, and the text content fonts in each text content area are consistent, then the text generation priority of each text content area is from top to bottom and from left to right.
[0070] If the text content areas are distributed and arranged in a staggered manner within the area to be collected, or the text content fonts of the text content areas are inconsistent, the center position of the area to be collected is determined, and the distance between each text area and the middle position of the area to be collected is judged to obtain the center distance value of each text content area. The text generation priority of each text content area is determined based on the center distance value of each text content area, wherein the larger the center distance value of the text content area, the higher the text generation priority of the text content area.
[0071] Step S323: Based on the text generation priorities of each text content area, perform preliminary text generation on the text feature extraction information of each text content area to obtain a preliminary text generation result. The purpose of performing preliminary text generation in sequence according to the text generation priorities is to improve the viewing efficiency of the user for the preliminarily generated text. In practical applications, when the user is eagerly waiting for text generation, by performing preliminary text generation in sequence according to the text generation priorities, the user can preferentially view key content, saving the user's waiting time.
[0072] Refer to Figure 1 and Figure 3 , Step S4: According to the image preprocessing information and the environmental voice preprocessing information, sequentially determine whether it is necessary to adjust the preliminary text generation result. If adjustment is required, perform the adjustment to obtain a text generation environment adjustment result. Step S4 specifically includes:
[0073] Step S41: Perform a first type of adjustment on the preliminary text generation result according to the relevant environmental image preprocessing information to obtain a first type of text generation adjustment result.
[0074] Specifically, based on image enhancement technology and image distortion correction algorithms, further process the relevant environmental image preprocessing information to obtain relevant environmental area reprocessed image information.
[0075] For the text generated for the text content area at the edge of the area to be collected in the preliminary text generation result, combine the relevant environmental area reprocessed image information to determine whether an error occurs. If no error occurs, obtain the first type of text generation adjustment result. At this time, the first type of text generation adjustment result is the preliminary text generation result. If an error occurs, comprehensively correct the preliminary text generation result according to the relevant environmental area reprocessed image information to obtain the first type of text generation adjustment result.
[0076] Now, an example is given for elaboration: If there is a "Spoon" in the text content area at the edge of the area to be collected, the text content generated in the preliminary text generation result is "Spoon". Combining the relevant environmental area reprocessed image information, it can be seen that there is a "White" on the left side of the "Spoon", and the spacing between the "White" and the "Spoon" is inconsistent with the spacing between other fonts in this text content area, being closer than the spacing between other fonts. Also, considering the coherence of the sentences in the text content area, when it is "of" here, the sentence is more coherent. Then, correct the "Spoon" at this position in the preliminary text generation result to "of" to obtain the first type of text generation adjustment result.
[0077] Step S42: Perform a second type of adjustment on the first type of adjusted text generation result according to the environmental voice preprocessing information to obtain a second type of text generation adjustment result.
[0078] Specifically, based on the artificial intelligence algorithm, the preprocessed environmental voice information is converted into the text content generated from the environmental voice.
[0079] Match the text content generated from the environmental voice with the generation result of the first type of adjusted text. Determine whether they match. If they match, no adjustment is required, and the generation result of the second type of adjusted text is obtained. At this time, the generation result of the second type of adjusted text is the same as that of the first type of adjusted text. If they do not match, mark the inconsistent content between the text content generated from the environmental voice and the generation result of the first type of adjusted text as the suspected content. Combine the image where the suspected content is located and the text around the suspected content to determine whether the suspected content needs to be adjusted, and then obtain the generation result of the second type of adjusted text.
[0080] Now, an example is given to illustrate: If the font in the text content area is a handwritten font, the accuracy of the generated generation result of the first type of adjusted text is relatively low. For example, in common handwritten fonts, "yaotiao" is prone to errors. Sometimes, in handwritten fonts, it may look like "qiong gui". When the generation result of the first type of adjusted text identifies it as "qiong gui", mark this place as the suspected content. According to the text content generated from the environmental voice, judge whether the suspected content is "yaotiao" based on the discussion among the people in the environment. At the same time, combine the sentences before and after the suspected content in the generation result of the first type of adjusted text to judge whether it is "yaotiao" or "qiong gui" here. If correction is needed, make the correction and adjustment to obtain the generation result of the second type of adjusted text.
[0081] Step S43, where the generation result of the first type of adjusted text and the generation result of the second type of adjusted text are combined to form the text generation environment adjustment result.
[0082] In actual application, first, adjust the preliminary text generation result through the relevant environmental image preprocessing information to obtain the generation result of the first type of adjusted text. Through the image of the area to be collected and the images of the relevant environmental areas, judge whether there are errors in the preliminary text generation result. If there are errors, make adjustments to improve the accuracy of text generation based on image recognition. Then, adjust the generation result of the first type of adjusted text through the preprocessed environmental voice information to obtain the generation result of the second type of adjusted text. Through the environmental voice during image acquisition, judge whether there are errors in the generation result of the first type of adjusted text. If there are errors, make adjustments to further improve the accuracy of text generation based on image recognition.
[0083] Refer to Figure 1 , step S5, according to the original image information and the original environmental voice information, analyze and identify the scene during the information collection of the area to be collected to obtain the collection scene recognition information. Step S5 specifically includes:
[0084] Step S51, obtaining a scene type data pool based on big data technology, wherein the scene type data pool includes user original type, brand marketing type, knowledge popularization type, news information type, and entertainment interaction type.
[0085] Step S52: Match the original image information and the original environmental voice information with the scene type data pool, determine the scene when the information is collected in the area to be collected, and obtain collection scene identification information.
[0086] Reference Figure 1 and Figure 4 In step S6, based on the scene recognition information, it is determined whether the scene at the time of image acquisition matches the scene of the text generation environment adjustment result, and whether adjustment is required. If adjustment is required, the text playback speed and the image playback speed are adjusted until the scenes match, and the final text generation result is output. Step S6 specifically includes:
[0087] Step S61 , matching the text generation environment adjustment result with the scene type data pool, determining the scene to which the generated text is suitable, and obtaining text generation scene identification information.
[0088] In step S62, the text generation scene recognition information is matched with the acquisition scene recognition information. If the text generation scene recognition information matches the acquisition scene recognition information, no adjustment is required and the final text generation result is output. In this case, the final text generation result is the text generation environment adjustment result. If the text generation scene recognition information does not match the acquisition scene recognition information, the playback speed of the generated text is adjusted. When the playback speed of the generated text reaches a preset text playback speed threshold, the adjustment of the generated text playback speed is stopped and the image playback speed is restarted. This process continues until the text generation scene recognition information matches the acquisition scene recognition information when each image is played, and the final text generation result is output.
[0089] In actual application, when subtitles are automatically generated for user videos of the user original type, the generated text often does not correspond to the image in the video. In the embodiment of the present application, the generated text playback speed and the image playback speed are adjusted to achieve accurate matching of the scene corresponding to the generated text and the scene corresponding to the image. Among them, the purpose of controlling the generated text playback speed within the preset text playback speed threshold is to prevent the text playback speed from being too fast or too slow, affecting the user's viewing experience.
[0090] In actual application, the text generation scene recognition information is obtained by matching the collected scene recognition information with the text generation scene recognition information. When the image of the area to be collected is played, the scene corresponding to the collected information is consistent with the scene function matching corresponding to the generated text, which facilitates the accurate matching of the generated text with the corresponding image and voice, and facilitates a good viewing experience for users.
[0091] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.
Claims
1. A text generation method based on image recognition, characterized in that: include: Step S1, obtaining a to-be-collected area, determining a related environmental area based on the to-be-collected area, collecting information from the to-be-collected area and the related environmental area to obtain original collected information, wherein the original collected information includes original image information and original environmental voice information; Step S2, preprocessing the original image information to obtain image preprocessing information, and preprocessing the original environmental voice information to obtain environmental voice preprocessing information; Step S3, reprocessing the preprocessed image information to obtain reprocessed image information of the acquisition area, and generating a preliminary text generation result based on the reprocessed image information of the acquisition area; Step S4, judging whether the preliminary text generation result needs to be adjusted based on the image preprocessing information and the environmental voice preprocessing information, and if so, performing the adjustment to obtain a text generation environment adjustment result; Step S5, analyzing and identifying the scene when collecting information in the collection area based on the original image information and the original environmental voice information to obtain collection scene identification information; Step S6, based on the acquisition scene recognition information, determines whether the scene during image acquisition matches the scene of the text generation environment adjustment result, and determines whether adjustment is needed. If adjustment is needed, adjust the text playback speed and image playback speed until the scenes match, and output the final text generation result.
2. The method for generating text based on image recognition according to claim 1, characterized in that: Step S1 specifically includes: Determine the relevant environmental area based on the area to be collected and the scene complexity; The original image information is obtained by performing image acquisition operation based on a high-definition camera, wherein the original image information includes image information of the acquisition area and image information of the related environmental area; When performing an image acquisition operation, the ambient sound around the acquisition point is synchronously acquired to obtain original ambient voice information, and the original image information at the same time is matched with the original ambient voice information based on the acquisition time.
3. The method for generating text based on image recognition according to claim 2, characterized in that: Step S2 specifically includes: Performing a preprocessing operation on the acquisition area image information and the related environment area image information in the original image information based on the image noise reduction technology to obtain image preprocessing information, wherein the image preprocessing information includes the acquisition area preprocessing image information and the related environment preprocessing image information; Based on the speech segmentation technology, the original environmental speech information is segmented to obtain a speech logical segment dataset, where the speech logical segment dataset includes multiple speech logical segments. Based on the time of each speech logical segment, whether the speech logical segments overlap is determined. If overlap, the timbre and content fluency of each speech logical segment are used to determine whether it is mis-segmented. If it is mis-segmented, the speech logical segment dataset is adjusted and updated; Determine whether there is human voice in each voice logic segment, filter out the voice logic segments without human voice and delete them, filter out the voice logic segments with human voice and mark them as valid voice logic segments, and pre-process the valid voice logic segments based on pre-emphasis processing technology and windowing processing technology to obtain environmental voice preprocessing information.
4. The method for generating text based on image recognition according to claim 3, wherein: Step S3 specifically includes: Based on image enhancement technology and image distortion correction algorithm, the pre-processed image information of the acquisition area is further processed to obtain the re-processed image information of the acquisition area; The image information is processed based on the collected area to initially generate the preliminary text generation result.
5. The method for generating text based on image recognition according to claim 4, characterized in that: The steps of reprocessing the image information based on the collected area to generate a preliminary text generation result specifically include: Marking the area containing text in the reprocessed image information of the acquisition area as a text content area, and extracting features of the text content in each text content area to obtain a plurality of text feature extraction information; Based on the distribution and arrangement of each text content area in the area to be collected and the font, the importance of the text content in each text content area is determined to obtain the text generation priority of each text content area; Based on the text generation priority of each text content area, preliminary text generation is performed on the text feature extraction information of each text content area in turn to obtain a preliminary text generation result.
6. The method for generating text based on image recognition according to claim 5, characterized in that: Step S4 specifically includes: Performing a first type of adjustment on the initial text generation result according to relevant environmental image preprocessing information to obtain a first type of text generation adjustment result; Performing a second type of adjustment on the first type of adjusted text generation result according to the environmental speech preprocessing information to obtain a second type of text generation adjustment result; Among them, the first type of text generation adjustment results and the second type of text generation adjustment results are combined to form the text generation environment adjustment results.
7. The method for generating text based on image recognition according to claim 6, characterized in that: Step S5 specifically includes: Based on big data technology, the scenario type data pool is obtained, where the scenario type data pool includes user original type, brand marketing type, knowledge popularization type, news information type, and entertainment interaction type; The original image information and the original environmental voice information are matched with the scene type data pool to determine the scene when the information is collected in the area to be collected to obtain the collection scene identification information.
8. The method for generating text based on image recognition according to claim 7, characterized in that: Step S6 specifically includes: Matching the text generation environment adjustment result with the scene type data pool to determine the scene suitable for the generated text and obtain text generation scene recognition information; Match the text generation scene recognition information with the collection scene recognition information. If the text generation scene recognition information matches the collection scene recognition information, there is no need to adjust the output of the final text generation result. If the text generation scene recognition information does not match the collection scene recognition information, adjust the generated text playback speed. When the generated text playback speed is adjusted to the preset text playback speed threshold, stop adjusting the generated text playback speed and start adjusting the image playback speed until the text generation scene recognition information matches the collection scene recognition information during playback, and output the final text generation result.