• inari@piefed.zip
    link
    fedilink
    English
    arrow-up
    35
    ·
    2 days ago

    I just automatically assume Google trains their models on absolutely all of the data is has access to

  • Elvith Ma'for@feddit.org
    link
    fedilink
    English
    arrow-up
    33
    arrow-down
    2
    ·
    2 days ago

    also seemed to be able to regurgitate accurate information related to unreleased mechanics that the developer says they had only written into GDocs the day before the player’s AI interaction.

    Then I don’t believe the “training” argument. Training takes time. They don’t release a new model daily.

    Either it had access to this doc while it was researching for the answer (that might be accidentally be some missing access controls or some link was shared somewhere with that user or…) or might have had a good guess. Depending on the game, players will be speculating about new patches, how to fix/change certain mechanics and such and the AI might have found/trained on this and just made a lucky guess. Same for the character name - if you know the other names and maybe there’s an (unconscious) pattern on how they’re named, it could also be a lucky guess.

    • AuroraZzz@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      ·
      2 days ago

      The ai was probably using the Google docs embedded into a vectordb. No training involved and a very common technique to do

      • Elvith Ma'for@feddit.org
        link
        fedilink
        English
        arrow-up
        2
        ·
        2 days ago

        Yeah, but then there should be access controls about which vectorization should be accessible for queries by whom…

    • fodor@lemmy.zip
      link
      fedilink
      English
      arrow-up
      2
      arrow-down
      9
      ·
      2 days ago

      That’s a definitional question. What counts as “training”? Is “search” different from training? Sometimes, but it depends what you mean.

      • yellow [she/her]@lemmy.blahaj.zone
        link
        fedilink
        English
        arrow-up
        26
        ·
        2 days ago

        “Training” refers to a process where training data is passed backwards through a model in order to modify the model’s weights. This happens before the model is deployed to the public. Search happens at (or right before) inference time (i.e. when the model is actually used) and does not modify the model weights.

      • Miaou@jlai.lu
        link
        fedilink
        English
        arrow-up
        4
        arrow-down
        1
        ·
        2 days ago

        Of course they are different. You can guess because we even have two different words for it

  • Leon@pawb.social
    link
    fedilink
    English
    arrow-up
    12
    ·
    2 days ago

    Google says Gemini can access Google Docs information, but only does so when given express permission

    Gemini is the trained model. The allegations isn’t that Gemini is snooping, it’s that Google stole the data to train the model. This is a non answer.

    The real answer is of course Google trained their models on anything you have stored with them. You’d be awfully naïve to think otherwise.

  • Blackmist@feddit.uk
    link
    fedilink
    English
    arrow-up
    7
    ·
    2 days ago

    Well I just tested it with something it should know if it was doing that (some API documentation we don’t publicly release), and it just made up some complete nonsense.

    So… Maybe not?

    • Nurse_Robot@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      ·
      2 days ago

      If it does train, that doesn’t mean it will use that data immediately or consistently reference one data point. I don’t think your test proves or disproves anything

      • Blackmist@feddit.uk
        link
        fedilink
        English
        arrow-up
        4
        ·
        2 days ago

        Probably. The document is specific to my software, and has been there a long time. Long enough that it should be in the training data, although getting out to pull that particular bit out could be a pain.

        Maybe they only trained on documents with certain privacy options set, or use some heuristics to determine if they should use it or not. Maybe the person in the article had his data hacked and shared on a fan site.

        Hard to tell, and it’s not like Google are going to tell us either way.

  • 𝙈𝙞𝙖@quokk.au
    link
    fedilink
    English
    arrow-up
    7
    ·
    2 days ago

    It also seemed to be able to regurgitate accurate information related to unreleased mechanics that the developer says they had only written into GDocs the day before the player’s AI interaction.

    Is training down daily now? I thought it was still a huge process, hence new models coming out like once a year or so.

    • tiramichu@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      8
      ·
      2 days ago

      Exactly.

      This is basically complete confirmation that the source of the data was NOT via ‘training’ because training is not that fast.

      Now, I’m not saying Google don’t train their AI on gdocs. It wouldn’t surprise me at all.

      But in this specific case, training cannot be where that information came from. There must be some other route of context-based access to the info.

      • Sendaris@lemmy.world
        link
        fedilink
        English
        arrow-up
        7
        ·
        2 days ago

        Seems more likely the Dev has accidentally set sharing permissions on their GDocs or Google Drive that allowed it to be indexed.

        • Dave.@aussie.zone
          link
          fedilink
          English
          arrow-up
          2
          ·
          2 days ago

          It’s quite possible that permission was accidentally given somewhere in Google’s platform when it was begging them to allow Gemini access “to improve your Google experience”.

          A phone pop-up, a notification in Gmail, Google Home, you name it, there’s plenty of places in the Google ecosystem to put a prompt in that’s worded in just the right way so that you think you’re sharing a slice and Google’s taking the whole pie.

  • Azzu@leminal.space
    link
    fedilink
    English
    arrow-up
    3
    arrow-down
    1
    ·
    2 days ago

    I wouldn’t be surprised if this was an attempt at marketing the game, lol