<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD with OASIS Tables with MathML3 v1.4 20220324//EN" "JATS-archive-oasis-tables-mathml3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" dtd-version="1.4" article-type="research-article">
  <front>
    <journal-meta>
      <journal-id journal-id-type="eissn">2588-0101</journal-id>
      <journal-title-group>
        <journal-title xml:lang="ru">Вестник Евразийской науки</journal-title>
        <journal-title xml:lang="en">The Eurasian Scientific Journal</journal-title>
      </journal-title-group>
      <publisher>
        <publisher-name>ООО «Издательство «Мир науки»</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="uri">https://esj.today/55NZVN623.html</article-id>
      <title-group>
        <article-title xml:lang="ru">Восстановление пропущенных значений в данных гидрометеорологических наблюдений с использованием машинного обучения (на примере реки Белая,
Республика Башкортостан)</article-title>
        <trans-title-group xml:lang="en">
          <trans-title>Missing values recovering in hydrometeorological data using machine learning (a case study from the Belaya river, republic of Bashkortostan)</trans-title>
        </trans-title-group>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name name-style="eastern">
            <surname>Тараканов</surname>
            <given-names>Денис Анатольевич</given-names>
          </name>
          <name-alternatives>
            <name xml:lang="ru" name-style="eastern">
              <surname>Тараканов</surname>
              <given-names>Денис Анатольевич</given-names>
            </name>
            <name xml:lang="en" name-style="western">
              <surname>Tarakanov</surname>
              <given-names>Denis Anatolyevich</given-names>
            </name>
          </name-alternatives>
          <email>tarakanov021098@gmail.com</email>
          <contrib-id contrib-id-type="orcid">https://orcid.org/0000-0003-0253-8624</contrib-id>
          <xref ref-type="aff" rid="aff1"/>
        </contrib>
        <aff-alternatives id="aff1">
          <aff>
            <institution xml:lang="ru">ФГАОУ ВО «ФГБОУ ВО «Уфимский университет науки и технологий» (Уфа, Россия)</institution>
          </aff>
          <aff>
            <institution xml:lang="en">Ufa University of Science and Technology (Ufa, Russia)</institution>
          </aff>
        </aff-alternatives>
      </contrib-group>
      <pub-date pub-type="epub" iso-8601-date="2024-02-12">
        <day>12</day>
        <month>02</month>
        <year>2024</year>
      </pub-date>
      <pub-date date-type="collection">
        <year>2023</year>
      </pub-date>
      <volume>15</volume>
      <issue>6</issue>
      <elocation-id>55NZVN623</elocation-id>
      <history>
        <date date-type="received" iso-8601-date="2023-12-15">
          <day>15</day>
          <month>12</month>
          <year>2023</year>
        </date>
        <date date-type="accepted" iso-8601-date="2024-02-05">
          <day>05</day>
          <month>02</month>
          <year>2024</year>
        </date>
      </history>
      <permissions>
        <copyright-year>2023</copyright-year>
        <copyright-holder xml:lang="ru">Тараканов Денис Анатольевич</copyright-holder>
        <copyright-holder xml:lang="en">Tarakanov Denis Anatolyevich</copyright-holder>
        <license xlink:href="https://creativecommons.org/licenses/by/4.0/">
          <license-p>CC BY 4.0</license-p>
        </license>
      </permissions>
      <self-uri xlink:type="simple" xlink:href="https://esj.today/55NZVN623.html">https://esj.today/55NZVN623.html</self-uri>
      <abstract xml:lang="ru">
        <p>Поскольку решение проблемы с пропущенными данными имеет ключевое значение при исследованиях, связанных с пространственно-климатическими и гидрологическими изменениями, выявлением и оценкой трендов и закономерностей, в данной работе исследуется возможность использования машинного обучения, как инструмента для восстановления пропущенных значений во временных рядах данных гидрометеорологических наблюдений. Предложен алгоритм заполнения пропусков, предполагающий три этапа. На первом этапе определяются непрерывные наборы данных на смежных постах (временные ряды данных на постах наблюдений, расположенных вблизи с исходным постом наблюдения, не имеющие пропусков) с высокой корреляционной взаимосвязью. Второй этап заключается в формировании обучающего и тестируемого наборов данных: обучающий набор данных представляет собой все имеющиеся значения исходного ряда наблюдений и соответствующие им значения смежных постов наблюдения; тестируемый набор данных состоит из всех имеющихся пропущенных значений и соответствующих им значений со смежных постов. На третьем этапе осуществляется выбор подходящих методов машинного обучения и метрик для оценки точности и эффективности разработанных моделей восстановления пропусков. Используя предложенный подход, получены наиболее оптимальные модели, позволяющие с высокой точностью восстановить пропущенные значения: для временных рядов расхода воды такими моделями являются множественная линейная регрессия и многослойный персептрон; для уровня воды — многослойный персептрон; для количества осадков — многослойный персептрон; для температуры воздуха — множественная линейная регрессия. Следующим этапом работы является исследование возможности улучшения результатов работы рассмотренных моделей путем использования инструментов для определения оптимальных гиперпараметров.</p>
      </abstract>
      <trans-abstract xml:lang="en">
        <p>Since solving the problem with missing data is of key importance in studies related to spatial-temporal, climatic and hydrological changes, identifying and evaluating trends and patterns, this paper examines the possibility and effectiveness of using machine learning as a tool for restoring missing values in time series of hydrometeorological data. An algorithm for filling in gaps is proposed, which involves three stages. At the first stage, continuous data sets of hydrometeorological data at adjacent stations are determined (time series of data at stations located near the original station, which do not have gaps) with a high correlation relationship. The second stage consists in the formation of training and test datasets: the training dataset represents all available values of the initial station and the corresponding values of adjacent stations; The data set under test consists of all available missing values and their corresponding values from adjacent stations. At the third stage, appropriate machine learning methods and metrics are selected to evaluate the accuracy and effectiveness of the developed gaps recovery models. Using the proposed approach, the most optimal models were obtained, allowing to restore the missing values with high accuracy: for time series of streamflow, the optimal models are Multiple Linear Regression and Multilayer Perceptron; for time series of water level, the Multilayer Perceptron model; for total precipitation time series — Multilayer Perceptron; for air temperature data sets — Multiple Linear Regression. The next stage of the work is to investigate the possibility of improving the performance of the considered models by using tools to determine the optimal hyperparameters of the models.</p>
      </trans-abstract>
      <kwd-group xml:lang="ru">
        <title>Ключевые слова</title>
        <kwd>восстановление пропусков</kwd>
        <kwd>восстановление пропущенных значений</kwd>
        <kwd>гидрологические данные</kwd>
        <kwd>климатические данные</kwd>
        <kwd>машинное обучение</kwd>
        <kwd>река Белая</kwd>
      </kwd-group>
      <kwd-group xml:lang="en">
        <title>Keywords</title>
        <kwd>gaps-filling</kwd>
        <kwd>hydrological data</kwd>
        <kwd>climate data</kwd>
        <kwd>machine learning</kwd>
        <kwd>Belaya River</kwd>
      </kwd-group>
      <funding-group>
        <funding-statement xml:lang="ru">Исследование выполнено за счет гранта Российского научного фонда № 22-27-00598, https://rscf.ru/project/22-27-00598/</funding-statement>
      </funding-group>
    </article-meta>
  </front>
  <back>
    <ref-list>
      <ref id="ref1">
        <label>1</label>
        <mixed-citation xml:lang="ru">Stocker T. Climate Change 2013: The Physical Science Basis / Working Group I Contribution to the Fifth Assessment Report of the Intergovernmental Panel on Climate Change; Cambridge University Press. — 2014. 1535 pp.</mixed-citation>
      </ref>
      <ref id="ref2">
        <label>2</label>
        <mixed-citation xml:lang="ru">Cui T., Li Y., Yang L. et al. Non-monotonic changes in Asian Water Towers’ streamflow at increasing warming levels / Nature Communications. — 2023. — Vol. 14. — P. 1176.</mixed-citation>
      </ref>
      <ref id="ref3">
        <label>3</label>
        <mixed-citation xml:lang="ru">Jones A., Kuehnert J., Fraccaro P. et al. AI for climate impacts: applications in flood risk / Climate and Atmospheric Science. — 2023. — Vol. 6. — P. 63.</mixed-citation>
      </ref>
      <ref id="ref4">
        <label>4</label>
        <mixed-citation xml:lang="ru">Siabi N., Sanaeinejad S.H., Ghahraman B. Effective Method for Filling Gaps in Time Series of Environmental Remote Sensing Data: An Example on Evapotranspiration and Land Surface Temperature Images / Computers and Electronics in Agriculture. — 2022. — Vol. 193. — 106619.</mixed-citation>
      </ref>
      <ref id="ref5">
        <label>5</label>
        <mixed-citation xml:lang="ru">Park J., Müller J., Arora B. et al. Long-term missing value imputation for time series data using deep neural networks / Neural Computing and Applications. — 2023. — Vol. 35. — P. 9071–9091.</mixed-citation>
      </ref>
      <ref id="ref6">
        <label>6</label>
        <mixed-citation xml:lang="ru">Nafikova E., Aleksandrov D., Shaniyazova A., Bondar Ch. Hydroecological data recovery using artificial intelligence / E3S Web of Conferences: 2022 International Scientific and Practical Conference on Development and Modern Problems of Aquaculture, AQUACULTURE 2022. — 2023. Vol. 381. — 01036.</mixed-citation>
      </ref>
      <ref id="ref7">
        <label>7</label>
        <mixed-citation xml:lang="ru">Thi-Thu-Hong Phan. Machine Learning for Univariate Time Series Imputation / Conference: 2020 International Conference on Multimedia Analysis and Pattern Recognition (MAPR). — 2020. — 6 pp.</mixed-citation>
      </ref>
      <ref id="ref8">
        <label>8</label>
        <mixed-citation xml:lang="ru">Navada A., Ansari A.N., Patil S. and Sonkamble B.A. Overview of use of decision tree algorithms in machine learning / 2011 IEEE Control and System Graduate Research Colloquium, Shah Alam, Malaysia. — 2011. — P. 37–42.</mixed-citation>
      </ref>
      <ref id="ref9">
        <label>9</label>
        <mixed-citation xml:lang="ru">Breiman L. Random Forests / Machine Learning. — 2001. — Vol. 45. — P. 5–32.</mixed-citation>
      </ref>
      <ref id="ref10">
        <label>10</label>
        <mixed-citation xml:lang="ru">Chen T., Guestrin C. Xgboost: A scalable tree boosting system / In Proceedings of the 22nd Acm Sigkdd International Conference on Knowledge Discovery and Data Mining. — 2016. — P. 785–794.</mixed-citation>
      </ref>
      <ref id="ref11">
        <label>11</label>
        <mixed-citation xml:lang="ru">Haykin S.S. Neural Networks: A Comprehensive Foundation / Prentice Hall PTR. — 1999. — P. 842.</mixed-citation>
      </ref>
      <ref id="ref12">
        <label>12</label>
        <mixed-citation xml:lang="ru">Evelyn Fix, Hodges J.L.Jr. Discriminatory Analysis. Nonparametric Discrimination: Consistency Properties / International Statistical Review. — 1989. — Vol. 57. — P. 238–247.</mixed-citation>
      </ref>
      <ref id="ref13">
        <label>13</label>
        <mixed-citation xml:lang="ru">Armstrong, J.S., Collopy, F. Error measures for generalizing about forecasting methods: Empirical comparisons / International Journal of Forecasting. — 1992. — Vol. 8. — P. 69–80.</mixed-citation>
      </ref>
      <ref id="ref14">
        <label>14</label>
        <mixed-citation xml:lang="ru">Mean Squared Error / The Concise Encyclopedia of Statistics. — 2008.</mixed-citation>
      </ref>
      <ref id="ref15">
        <label>15</label>
        <mixed-citation xml:lang="ru">Willmott CJ., Matsuura K. Advantages of the Mean Absolute Error (MAE) over the Root Mean Square Error (RMSE) in Assessing Average Model Performance. — 2005. — Vol. 30. — P. 79–82.</mixed-citation>
      </ref>
      <ref id="ref16">
        <label>16</label>
        <mixed-citation xml:lang="ru">Pedregosa et al. Scikit-learn: Machine Learning in Python / JMLR 12. — 2011. — P. 2825–2830.</mixed-citation>
      </ref>
      <ref id="ref17">
        <label>17</label>
        <mixed-citation xml:lang="ru">Hunter J.D. Matplotlib: A 2D Graphics Environment / Computing in Science &amp; Engineering. — 2007. — Vol. 9. — P. 90–95.</mixed-citation>
      </ref>
      <ref id="ref18">
        <label>18</label>
        <mixed-citation xml:lang="ru">Tarakanov Denis Anatolyevich</mixed-citation>
      </ref>
      <ref id="ref19">
        <label>19</label>
        <mixed-citation xml:lang="ru">Ufa University of Science and Technology, Ufa, Russia</mixed-citation>
      </ref>
      <ref id="ref20">
        <label>20</label>
        <mixed-citation xml:lang="ru">E-mail: tarakanov021098@gmail.com</mixed-citation>
      </ref>
      <ref id="ref21">
        <label>21</label>
        <mixed-citation xml:lang="ru">ORCID: https://orcid.org/0000-0003-0253-8624</mixed-citation>
      </ref>
      <ref id="ref22">
        <label>22</label>
        <mixed-citation xml:lang="ru">RSCI: https://elibrary.ru/author_profile.asp?id=1153499</mixed-citation>
      </ref>
      <ref id="ref23">
        <label>23</label>
        <mixed-citation xml:lang="ru">SCOPUS: https://www.scopus.com/authid/detail.url?authorId=57218675021</mixed-citation>
      </ref>
      <ref id="ref24">
        <label>24</label>
        <mixed-citation xml:lang="ru">Missing values recovering in hydrometeorological
data using machine learning (a case study from the Belaya river, republic of Bashkortostan)</mixed-citation>
      </ref>
      <ref id="ref25">
        <label>25</label>
        <mixed-citation xml:lang="ru">Abstract. Since solving the problem with missing data is of key importance in studies related to spatial-temporal, climatic and hydrological changes, identifying and evaluating trends and patterns, this paper examines the possibility and effectiveness of using machine learning as a tool for restoring missing values in time series of hydrometeorological data. An algorithm for filling in gaps is proposed, which involves three stages. At the first stage, continuous data sets of hydrometeorological data at adjacent stations are determined (time series of data at stations located near the original station, which do not have gaps) with a high correlation relationship. The second stage consists in the formation of training and test datasets: the training dataset represents all available values of the initial station and the corresponding values of adjacent stations; The data set under test consists of all available missing values and their corresponding values from adjacent stations. At the third stage, appropriate machine learning methods and metrics are selected to evaluate the accuracy and effectiveness of the developed gaps recovery models. Using the proposed approach, the most optimal models were obtained, allowing to restore the missing values with high accuracy: for time series of streamflow, the optimal models are Multiple Linear Regression and Multilayer Perceptron; for time series of water level, the Multilayer Perceptron model; for total precipitation time series — Multilayer Perceptron; for air temperature data sets — Multiple Linear Regression. The next stage of the work is to investigate the possibility of improving the performance of the considered models by using tools to determine the optimal hyperparameters of the models.</mixed-citation>
      </ref>
      <ref id="ref26">
        <label>26</label>
        <mixed-citation xml:lang="ru">Keywords: gaps-filling; hydrological data; climate data; machine learning; Belaya River</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>
