TL;DR: In the last year, the Wikimedia Foundation has fired several union organizers, including those that worked on the Community Tech team - a team dedicated to building features for the volunteer community that edits Wikipedia.

As the Wiki Workers Union tries to get the Wikimedia Foundation to recognize their union, it is worth remembering that this is not the first time that the Foundation has worked against the community.

Wikimedia Enterprise is a betrayal of the volunteer movement community of Wikipedia editors, as the Wikimedia Foundation is providing privileged access to big tech AI companies to the Wikipedia corpus - a body of work that the Foundation does not own.

Movement volunteer communities contributed to Wikipedia under copyleft licenses - licenses that work to ensure that the work remains free (as in speech). The big tech AI companies do not license derivative works under copyleft licenses and often do not even attribute where the works came from.

This means that volunteers are working for big tech for free, and the Wikimedia Foundation is selling privileged access to that free labor.

It is against that backdrop that the current unionization struggle unfolds.

  • yoasif@fedia.ioOP
    link
    fedilink
    arrow-up
    1
    ·
    5 days ago

    if a C&D doesn’t accuse you of anything illegal, then it absolutely cannot compel you under any circumstance:

    This doesn’t have to mean criminal penalties. If WMF simply tells the scrapers that they are no longer authorized to access their systems, they can litigate against companies who continue to breach their request to discontinue scraping. That can be a civil action.

    I can’t imagine that Wikimedia is somehow required to feed the LLMs, and that they cannot simply request that they stop. You and I may not get a lot of traction there, but $263M definitely ought to make it possible to hire a lawyer to defend against (now) unauthorized access.

    We don’t know that because WMF hasn’t bothered to try defending contributors.

    Very false. Bot detection and blocking efforts have always been pretty documented

    I don’t see how bot detection shows that WMF is trying to defend contributors’ IP.

    Instead, they created a glide path for the pirates taking advantage of them.

    Before Wikimedia Enterprise it was way easier. Again, API Access was, until very recently, unlimited and free. I’m not sure how Enterprise is the glide path here.

    Before Wikimedia Enterprise they weren’t getting paid and the LLM companies may have had to depend on residential proxies – if the bot detection you mention was successful, for example. The glide path allows Enterprise customers to not experience rate limits or rely on shoddy infrastructure. Clearly, it’s better to not be in the shadows, if you are trying to evade detection.

    PS: Were they using the API or were they scraping? If they were using the API, why couldn’t WMF just revoke their keys?

    • Aatube@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      1
      ·
      5 days ago

      Fair point on using ToS.

      Before Wikimedia Enterprise they weren’t getting paid and the LLM companies may have had to depend on residential proxies

      It is extremely cheap to rely on residential proxies and certainly magnitudes cheaper than Enterprise. The free infrastructure is absolutely not shoddy and I’m not sure why you think it is, though making all scrapers use it would probably make it shoddy.

      If they were using the API, why couldn’t WMF just revoke their keys?

      The API has no keys. It’s commons public access. If you want to revoke access, you’d have to identify the IPs and manually block them.

      You’re also still assuming this is violating intellectual property. Customers of Wikimedia Enterprise are forced to sign Terms of Use (accessible at https://dashboard.enterprise.wikimedia.com/signup/) that legally bind them to respect the license of the content even though, as you point out, they are already bound to respect the license. To maintain access to Enterprise, you must not violate that license, and with that the WMF can send DMCA C&Ds.

      Since Enterprise does have API keys, it in fact gives WMF control over exploitation of it enforceable in the way you mention. Enterprise is what defends our copyright.

      • yoasif@fedia.ioOP
        link
        fedilink
        arrow-up
        1
        ·
        4 days ago

        Enterprise is what defends our copyright.

        But it clearly isn’t. Hence my mention of it in my post and why I see it as a betrayal.

        • Aatube@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          0
          ·
          4 days ago

          would you kindly respond to what I said :)? restating a conclusion i’ve responded to is unfortunately not much for me to go off of

          • yoasif@fedia.ioOP
            link
            fedilink
            arrow-up
            1
            ·
            4 days ago

            Yes, I am assuming that this violates the Wikipedia IP, I explained that in my post. I’m not going to re-litigate that since you haven’t shown how I am wrong about that.

            Given that, a C&D for discontinuing scraping, followed by a suit to desist any further scraping would seem to be in order.

            • Aatube@lemmy.dbzer0.com
              link
              fedilink
              English
              arrow-up
              1
              ·
              2 days ago

              Enterprise does not violate the copyright (please say that instead of “intellectual property”: https://www.gnu.org/philosophy/not-ipr.html “These laws originated separately, evolved differently, cover different activities, have different rules, and raise different public policy issues.”). Customers of Wikimedia Enterprise are forced to sign Terms of Use (accessible at https://dashboard.enterprise.wikimedia.com/signup/) that legally bind them to respect the license of the content even though, as you point out, they are already bound to respect the license. To maintain access to Enterprise, you must not violate that license, and with that the WMF can send DMCA C&Ds.

              It’s also believed that Enterprise customers are not violating the license. For the LLM usecase:

              Overall, it is more likely than not if current precedent holds that training systems on copyrighted data will be covered by fair use in the United States, but there is significant uncertainty at time of writing.

              https://meta.wikimedia.org/wiki/Wikilegal/Copyright_Analysis_of_ChatGPT

              Fair use is a doctrine in United States law that permits limited use of copyrighted material without permission from the copyright holder.

              • yoasif@fedia.ioOP
                link
                fedilink
                arrow-up
                1
                ·
                2 days ago

                “More than likely” is a convenient analysis for Wikimedia Enterprise, I don’t agree with that - nor do I agree with Wikimedia not defending copyrights.

                I don’t think we’re going to get anywhere since you seem to think that the “more than likely” analysis is correct, while I don’t agree with that (nor have we seen the courts strike down the copyleft portions of the licenses as unworkable given LLM based “fair use”).

                Happy to move on here.

                To maintain access to Enterprise, you must not violate that license, and with that the WMF can send DMCA C&Ds.

                We’re really on the same page - you just need to take one step to realize that if the LLM vendors cannot or will not reciprocally license derivative works, they are clearly not relying on the license to copy the works.

                If instead, they are relying on fair use, it isn’t clear to me that they are not violating the license - since the license requires reciprocal licensing of derivative works.

                If they are relying on fair use, they are violating the license, and WMF can send C&Ds. Right?

                Where is the break in the logic?